Market Research
Applied market research for people who have to defend a number in a room. This
skill is about the operational craft: constructing a market size two independent
ways, reconciling the gap, cutting the market into segments that behave
differently, and fielding survey instruments that do not manufacture the answer
you hoped for.
When to use this skill
- Sizing a market for a board deck, investor memo, or funding request where
the number will be challenged line by line
- Reconciling a TAM you inherited — an analyst report says $12B, your
bottom-up build says $700M, and you need to explain the gap
- Segmenting a market before a pricing, packaging, or GTM decision
- Triangulating demand signals (search volume, inbound, win rates, analyst
data, competitor headcount) into one directional read
- Designing a survey to answer a market question — willingness to pay,
category awareness, switching intent — without leading the respondent
- Auditing someone else's sizing before you sign off on it
Inputs the skill expects
- The market definition in one sentence — including geography and buyer
- A top-down anchor (published market value) with its source and vintage
- Bottom-up unit economics — unit count, qualified share, annual value per unit
- The decision the number is feeding (investment size, hiring plan, pricing)
- Time horizon for SOM (1 year vs 3 years changes it by an order of magnitude)
- For surveys: population size, target margin of error, mode (panel, list, intercept)
Clarify First
Before generating, confirm these inputs. If any is unknown or vague, ASK — do not assume:
Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
Workflows
Workflow 1 — Build and reconcile TAM/SAM/SOM
- Write the market definition sentence first. Everything downstream inherits it.
- Build the top-down chain: published market value, then named filters that
each cut it (geography, segment, buyer qualification), each with a retention
fraction and a stated justification.
- Build the bottom-up chain independently: unit count from a countable source,
qualified share, annual value per unit, reachable share, expected win rate.
- Run the builder. It computes both chains, reconciles them layer by layer, and
flags implausible ratios and divergence.
- Resolve every
fail before the number leaves your machine. A warn needs a
sentence in the memo, not a fix.
python3 research-ops/market-research/scripts/tam_sam_som_builder.py \
--input research-ops/market-research/assets/sample_market_model.json \
--format text
Workflow 2 — Triangulate demand signals
- Collect every observable demand signal you have — search volume, inbound
lead velocity, win rate by segment, analyst growth rates, competitor hiring,
category conference attendance.
- Score each for source independence and directional strength.
- Run the triangulator to get a weighted demand index and, more importantly,
the list of signals that contradict each other.
- Investigate contradictions before averaging them away. A conflicting signal
is usually a segmentation boundary you have not drawn yet.
python3 research-ops/market-research/scripts/demand_signal_triangulator.py \
--input research-ops/market-research/assets/sample_demand_signals.json \
--format text
Workflow 3 — Audit a survey instrument before fielding
- Draft the instrument with the market question stated at the top.
- Run the auditor. It checks each item for leading language, double-barrelled
phrasing, absolutes, unbalanced or over-long scales, and missing escape
options.
- Check the sample-size verdict — it computes required n from population,
target margin of error, and confidence level.
- Fix every
fail, then re-run. Field only on a clean run.
python3 research-ops/market-research/scripts/survey_instrument_auditor.py \
--input research-ops/market-research/assets/sample_survey.json \
--format text
Decision frameworks
Which sizing method for which situation
| Situation |
Method |
Why |
| Established category, published reports exist |
[PROVEN] Top-down anchored, bottom-up as a check |
The anchor is defensible; bottom-up catches definition drift |
| New category, no analyst coverage |
[PROVEN] Bottom-up only, stated as such |
A top-down number for a category that does not exist yet is fiction |
| Adjacent expansion from an existing product |
[RECOMMENDED] Bottom-up from your own funnel conversion |
Your observed win rates beat any external estimate |
| Regulated market with registries |
[PROVEN] Bottom-up from the registry count |
Counting licensed entities is the strongest unit base available |
| Consumer market, behaviour-driven |
[RECOMMENDED] Top-down plus survey-derived incidence |
Unit counts exist but qualification requires stated behaviour |
Plausibility thresholds
These are the ratios the builder enforces. They are heuristics, not laws — but
crossing one without an explanation in the memo is how sizing loses credibility.
| Ratio |
Healthy range |
Flag when |
| SAM / TAM |
5% – 40% |
Above 60% — you are claiming almost the whole market is addressable |
| SOM / SAM (3-year) |
1% – 10% |
Above 20% — implies category leadership inside the horizon |
| SOM / TAM |
0.1% – 5% |
Above 5% for a pre-scale company |
| Bottom-up vs top-down TAM |
Within 3x |
Above 3x warn, above 10x fail — the two builds are answering different questions |
Survey sample size at 95% confidence
Required n for a proportion estimate, finite population corrected. Use these as
a sanity check on the auditor's output.
| Population |
±10% MoE |
±5% MoE |
±3% MoE |
| 500 |
81 |
218 |
341 |
| 5,000 |
95 |
357 |
880 |
| 100,000 |
96 |
383 |
1,056 |
| 1,000,000+ |
97 |
385 |
1,066 |
The jump from ±10% to ±5% quadruples cost for a band most market decisions do
not need. [RECOMMENDED] Field at ±10% for directional category questions and
reserve ±5% for pricing and packaging decisions where the band drives the choice.
Anti-Patterns
The Inherited TAM
Mistake: Copying a market size from an analyst report or a competitor's deck
into your own memo, adjusting the geography, and presenting it as your build.
Why it happens: The number is already large and already sourced, and building
bottom-up takes two days you do not think you have.
Instead: Use the published figure as the top-down anchor only, and always
build the bottom-up chain alongside it. The reconciliation gap is the most
informative artifact of the whole exercise — it tells you exactly which
definition the report used and yours does not.
The Multiplication Fantasy
Mistake: SOM computed as "if we capture 1% of the TAM" with no mechanism
behind the 1%.
Why it happens: It sounds modest, so nobody challenges it, and it produces a
convenient number without requiring a channel model.
Instead: Build SOM from reachable units times expected win rate, where both
come from something observed — your funnel, a pilot, or a comparable. If you
cannot name the channel that reaches those units, you do not have a SOM.
The Stale Anchor
Mistake: A four-year-old market report used at face value in a current memo.
Why it happens: It was the best available source when someone first built the
model, and nobody re-checks a number that has been in the deck for a year.
Instead: Record the vintage of every anchor. If it is more than 18 months
old, apply an explicit growth bridge with a stated CAGR and show both the raw
and bridged figures. An unbridged stale anchor invites the reviewer to discount
everything downstream of it.
The Leading Instrument
Mistake: Asking "How valuable would an automated reporting feature be to
your team?" and reporting the enthusiasm as demand evidence.
Why it happens: The team already believes in the feature, and the question is
written by the person who wants it built.
Instead: Ask about the current behaviour and its cost — "How many hours last
month did your team spend building reports manually?" — and let the demand fall
out of the numbers. Run every instrument through the auditor before fielding;
leading items are cheap to fix pre-field and impossible to fix post-field.
Segments That Do Not Behave Differently
Mistake: Cutting the market by company size or geography because that data is
available, then finding every segment has the same conversion and the same ACV.
Why it happens: Firmographic fields are in the CRM; behavioural ones are not.
Instead: Segment on the variable that changes the buying decision — trigger
event, existing tooling, regulatory obligation, or team structure. A segmentation
is only useful if the segments have measurably different win rates or values.
Files
| File |
Purpose |
scripts/tam_sam_som_builder.py |
Builds top-down and bottom-up TAM/SAM/SOM, reconciles them, flags implausible ratios |
scripts/survey_instrument_auditor.py |
Checks survey items for leading language, scale problems, and computes required sample size |
scripts/demand_signal_triangulator.py |
Weights and triangulates demand signals; surfaces contradictions and source concentration |
references/market-sizing-methods.md |
Method selection, filter design, growth bridges, worked reconciliation examples |
references/survey-design-methodology.md |
Question construction, scale design, sampling frames, mode effects, field QA |
assets/market-sizing-memo-template.md |
The memo structure a sizing number ships in |
assets/sample_market_model.json |
Runnable input for the TAM/SAM/SOM builder |
assets/sample_survey.json |
Runnable input for the survey auditor |
assets/sample_demand_signals.json |
Runnable input for the demand triangulator |
1---2name: market-research3description: Market sizing and market structure work — TAM/SAM/SOM built top-down and bottom-up then reconciled, segmentation, demand triangulation, and survey design. Use when sizing a market, writing a sizing memo, or fielding a survey.4license: MIT + Commons Clause5---6
7# Market Research
8
9Applied market research for people who have to defend a number in a room. This
10skill is about the operational craft: constructing a market size two independent
11ways, reconciling the gap, cutting the market into segments that behave
12differently, and fielding survey instruments that do not manufacture the answer
13you hoped for.
14
15## When to use this skill
16
17- **Sizing a market for a board deck, investor memo, or funding request** where
18 the number will be challenged line by line
19- **Reconciling a TAM you inherited** — an analyst report says $12B, your
20 bottom-up build says $700M, and you need to explain the gap
21- **Segmenting a market** before a pricing, packaging, or GTM decision
22- **Triangulating demand signals** (search volume, inbound, win rates, analyst
23 data, competitor headcount) into one directional read
24- **Designing a survey** to answer a market question — willingness to pay,
25 category awareness, switching intent — without leading the respondent
26- **Auditing someone else's sizing** before you sign off on it
27
28## Inputs the skill expects
29
30- The market definition in one sentence — including geography and buyer
31- A top-down anchor (published market value) with its source and vintage
32- Bottom-up unit economics — unit count, qualified share, annual value per unit
33- The decision the number is feeding (investment size, hiring plan, pricing)
34- Time horizon for SOM (1 year vs 3 years changes it by an order of magnitude)
35- For surveys: population size, target margin of error, mode (panel, list, intercept)
36
37## Clarify First
38
39Before generating, confirm these inputs. If any is unknown or vague, ASK — do not assume:
40
41- [ ] **Market definition — what is in and what is out** — the single biggest driver of the number; "dental software" and "dental practice management software for multi-chair EU practices" differ by 20x
42- [ ] **The decision this sizing supports** — a fundraise tolerates a wide TAM; a hiring plan needs a defensible SOM
43- [ ] **Time horizon for SOM** — 12-month obtainable share and 3-year obtainable share are different artifacts
44- [ ] **Whether a published anchor exists and its vintage** — a 2022 report in a 2026 memo needs an explicit growth bridge
45
46Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
47
48## Workflows
49
50### Workflow 1 — Build and reconcile TAM/SAM/SOM
51
521. Write the market definition sentence first. Everything downstream inherits it.
532. Build the top-down chain: published market value, then named filters that
54 each cut it (geography, segment, buyer qualification), each with a retention
55 fraction and a stated justification.
563. Build the bottom-up chain independently: unit count from a countable source,
57 qualified share, annual value per unit, reachable share, expected win rate.
584. Run the builder. It computes both chains, reconciles them layer by layer, and
59 flags implausible ratios and divergence.
605. Resolve every `fail` before the number leaves your machine. A `warn` needs a
61 sentence in the memo, not a fix.
62
63```bash
64python3 research-ops/market-research/scripts/tam_sam_som_builder.py \
65 --input research-ops/market-research/assets/sample_market_model.json \
66 --format text
67```
68
69### Workflow 2 — Triangulate demand signals
70
711. Collect every observable demand signal you have — search volume, inbound
72 lead velocity, win rate by segment, analyst growth rates, competitor hiring,
73 category conference attendance.
742. Score each for source independence and directional strength.
753. Run the triangulator to get a weighted demand index and, more importantly,
76 the list of signals that contradict each other.
774. Investigate contradictions before averaging them away. A conflicting signal
78 is usually a segmentation boundary you have not drawn yet.
79
80```bash
81python3 research-ops/market-research/scripts/demand_signal_triangulator.py \
82 --input research-ops/market-research/assets/sample_demand_signals.json \
83 --format text
84```
85
86### Workflow 3 — Audit a survey instrument before fielding
87
881. Draft the instrument with the market question stated at the top.
892. Run the auditor. It checks each item for leading language, double-barrelled
90 phrasing, absolutes, unbalanced or over-long scales, and missing escape
91 options.
923. Check the sample-size verdict — it computes required n from population,
93 target margin of error, and confidence level.
944. Fix every `fail`, then re-run. Field only on a clean run.
95
96```bash
97python3 research-ops/market-research/scripts/survey_instrument_auditor.py \
98 --input research-ops/market-research/assets/sample_survey.json \
99 --format text
100```
101
102## Decision frameworks
103
104### Which sizing method for which situation
105
106| Situation | Method | Why |
107|-----------|--------|-----|
108| Established category, published reports exist | **[PROVEN]** Top-down anchored, bottom-up as a check | The anchor is defensible; bottom-up catches definition drift |
109| New category, no analyst coverage | **[PROVEN]** Bottom-up only, stated as such | A top-down number for a category that does not exist yet is fiction |
110| Adjacent expansion from an existing product | **[RECOMMENDED]** Bottom-up from your own funnel conversion | Your observed win rates beat any external estimate |
111| Regulated market with registries | **[PROVEN]** Bottom-up from the registry count | Counting licensed entities is the strongest unit base available |
112| Consumer market, behaviour-driven | **[RECOMMENDED]** Top-down plus survey-derived incidence | Unit counts exist but qualification requires stated behaviour |
113
114### Plausibility thresholds
115
116These are the ratios the builder enforces. They are heuristics, not laws — but
117crossing one without an explanation in the memo is how sizing loses credibility.
118
119| Ratio | Healthy range | Flag when |
120|-------|---------------|-----------|
121| SAM / TAM | 5% – 40% | Above 60% — you are claiming almost the whole market is addressable |
122| SOM / SAM (3-year) | 1% – 10% | Above 20% — implies category leadership inside the horizon |
123| SOM / TAM | 0.1% – 5% | Above 5% for a pre-scale company |
124| Bottom-up vs top-down TAM | Within 3x | Above 3x warn, above 10x fail — the two builds are answering different questions |
125
126### Survey sample size at 95% confidence
127
128Required n for a proportion estimate, finite population corrected. Use these as
129a sanity check on the auditor's output.
130
131| Population | ±10% MoE | ±5% MoE | ±3% MoE |
132|-----------|----------|---------|---------|
133| 500 | 81 | 218 | 341 |
134| 5,000 | 95 | 357 | 880 |
135| 100,000 | 96 | 383 | 1,056 |
136| 1,000,000+ | 97 | 385 | 1,066 |
137
138The jump from ±10% to ±5% quadruples cost for a band most market decisions do
139not need. **[RECOMMENDED]** Field at ±10% for directional category questions and
140reserve ±5% for pricing and packaging decisions where the band drives the choice.
141
142## Anti-Patterns
143
144### The Inherited TAM
145**Mistake:** Copying a market size from an analyst report or a competitor's deck
146into your own memo, adjusting the geography, and presenting it as your build.
147**Why it happens:** The number is already large and already sourced, and building
148bottom-up takes two days you do not think you have.
149**Instead:** Use the published figure as the top-down anchor only, and always
150build the bottom-up chain alongside it. The reconciliation gap is the most
151informative artifact of the whole exercise — it tells you exactly which
152definition the report used and yours does not.
153
154### The Multiplication Fantasy
155**Mistake:** SOM computed as "if we capture 1% of the TAM" with no mechanism
156behind the 1%.
157**Why it happens:** It sounds modest, so nobody challenges it, and it produces a
158convenient number without requiring a channel model.
159**Instead:** Build SOM from reachable units times expected win rate, where both
160come from something observed — your funnel, a pilot, or a comparable. If you
161cannot name the channel that reaches those units, you do not have a SOM.
162
163### The Stale Anchor
164**Mistake:** A four-year-old market report used at face value in a current memo.
165**Why it happens:** It was the best available source when someone first built the
166model, and nobody re-checks a number that has been in the deck for a year.
167**Instead:** Record the vintage of every anchor. If it is more than 18 months
168old, apply an explicit growth bridge with a stated CAGR and show both the raw
169and bridged figures. An unbridged stale anchor invites the reviewer to discount
170everything downstream of it.
171
172### The Leading Instrument
173**Mistake:** Asking "How valuable would an automated reporting feature be to
174your team?" and reporting the enthusiasm as demand evidence.
175**Why it happens:** The team already believes in the feature, and the question is
176written by the person who wants it built.
177**Instead:** Ask about the current behaviour and its cost — "How many hours last
178month did your team spend building reports manually?" — and let the demand fall
179out of the numbers. Run every instrument through the auditor before fielding;
180leading items are cheap to fix pre-field and impossible to fix post-field.
181
182### Segments That Do Not Behave Differently
183**Mistake:** Cutting the market by company size or geography because that data is
184available, then finding every segment has the same conversion and the same ACV.
185**Why it happens:** Firmographic fields are in the CRM; behavioural ones are not.
186**Instead:** Segment on the variable that changes the buying decision — trigger
187event, existing tooling, regulatory obligation, or team structure. A segmentation
188is only useful if the segments have measurably different win rates or values.
189
190## Files
191
192| File | Purpose |
193|------|---------|
194| `scripts/tam_sam_som_builder.py` | Builds top-down and bottom-up TAM/SAM/SOM, reconciles them, flags implausible ratios |
195| `scripts/survey_instrument_auditor.py` | Checks survey items for leading language, scale problems, and computes required sample size |
196| `scripts/demand_signal_triangulator.py` | Weights and triangulates demand signals; surfaces contradictions and source concentration |
197| `references/market-sizing-methods.md` | Method selection, filter design, growth bridges, worked reconciliation examples |
198| `references/survey-design-methodology.md` | Question construction, scale design, sampling frames, mode effects, field QA |
199| `assets/market-sizing-memo-template.md` | The memo structure a sizing number ships in |
200| `assets/sample_market_model.json` | Runnable input for the TAM/SAM/SOM builder |
201| `assets/sample_survey.json` | Runnable input for the survey auditor |
202| `assets/sample_demand_signals.json` | Runnable input for the demand triangulator |