Patient Population Sizer — The Epidemiology-to-TAM Funnel
The gap between "50 million Americans have disease X" and "how many patients will take this drug" is where most revenue forecasts go wrong. A headline prevalence number is not a market. This skill builds a rigorous waterfall from total disease burden down to the treatable addressable population for a specific therapeutic asset, with sourced estimates at every step. The output feeds directly into peak-sales-forecaster as the patient volume input.
Every biotech pitch deck inflates patient numbers. This skill deflates them to reality.
How to Run
Input
| Parameter |
Required? |
Example |
| Disease/indication |
Yes |
Non-small cell lung cancer (NSCLC) |
| Geography |
Yes |
US, EU5, Japan, Global |
| Line of therapy |
Recommended |
1L metastatic, 2L+, adjuvant |
| Biomarker requirement |
If applicable |
PD-L1 TPS >= 50%, EGFR-mutant |
| Modality constraints |
If applicable |
IV infusion only (limits to infusion centers) |
| Competitive context |
Recommended |
Current SOC, anticipated entrants |
Steps
Epidemiology data: For sourced prevalence/incidence/survival (SEER 2026, GLOBOCAN 2022), cardiometabolic/CNS/immunology base rates (CDC, GBD), and a worked line-of-therapy attrition funnel, use references/epidemiology-by-indication.md. It is the single source of truth for biomarker/disease prevalence in the suite. Note global GBD figures understate US/EU addressable populations.
Step 1 — Establish Disease Burden (Top of Funnel)
Start with the broadest defensible population estimate. Use the data source hierarchy — always cite the highest-tier source available:
| Tier |
Source |
Strengths |
Limitations |
| 1 |
GBD (Global Burden of Disease) |
Standardized global methodology, 204 countries |
Modeled estimates, 2-3yr lag |
| 2 |
SEER / national cancer registries |
Observed incidence, stage-specific |
US/country-specific, cancer-only |
| 3 |
Claims databases (Truven, Optum, IQVIA) |
Real-world treatment patterns |
Commercially insured bias, US-centric |
| 4 |
Published epidemiological studies |
Disease-specific depth |
Single population, may be outdated |
| 5 |
KOL estimates / company projections |
Forward-looking |
Unverifiable, potential bias |
Prevalence vs. incidence. Use prevalence for chronic diseases where patients accumulate (e.g., rheumatoid arthritis, diabetes). Use incidence for diseases with rapid turnover (e.g., metastatic cancers where median survival < 2 years, acute infections). For oncology, annual incidence is typically the right starting point; for immunology, point prevalence.
Step 2 — Apply Diagnostic Rate Filter
Not all patients with a disease are diagnosed. This filter is often the largest single reduction in the funnel.
| Disease Category |
Typical Diagnosis Rate |
Key Drivers |
| Common cancers (breast, lung, colon) |
85-95% |
Screening programs, symptomatic presentation |
| Rare cancers |
70-85% |
May be misdiagnosed; referral delays |
| Autoimmune diseases |
50-70% |
Lengthy diagnostic odyssey, overlapping symptoms |
| Rare genetic diseases |
5-30% |
Requires genetic testing, low awareness |
| Neurodegenerative (early-stage) |
30-60% |
Gradual onset, normalization of symptoms |
Formula: Diagnosed population = Disease burden x Diagnosis rate
Step 3 — Apply Treatment-Seeking Filter
Not all diagnosed patients seek or accept treatment.
| Factor |
Adjustment |
Rationale |
| Asymptomatic early-stage disease |
-20-40% |
Watchful waiting, physician and patient reluctance |
| Stigmatized conditions |
-15-30% |
Mental health, substance abuse, sexual health |
| Elderly/frail patients |
-10-20% |
May not be offered aggressive therapy |
| Highly symptomatic, life-threatening |
-5% or less |
Strong motivation to treat |
| Disease with clear treatment guidelines |
-5-10% |
Guideline-concordant care drives treatment |
Formula: Treatment-seeking = Diagnosed x Treatment-seeking rate
Step 4 — Apply Treatment Eligibility Filters
Clinical trial inclusion/exclusion criteria and label restrictions reduce the eligible population.
| Filter |
Typical Impact |
Examples |
| Performance status (ECOG) |
-10-20% |
ECOG 0-1 required; ECOG 2+ excluded |
| Organ function requirements |
-5-15% |
Adequate hepatic/renal function |
| Prior therapy requirements |
Variable |
Must have failed 1L; may exclude pretreated |
| Contraindications |
-5-10% |
Autoimmune history for IO, cardiac for certain TKIs |
| Age restrictions |
-2-5% |
Pediatric exclusion, upper age limits |
| Comorbidities |
-10-20% |
Uncontrolled diabetes, active infections |
Formula: Eligible = Treatment-seeking x (1 - sum of exclusion rates)
Step 5 — Apply Biomarker Prevalence Filter
If the drug requires biomarker selection, this can dramatically reduce the addressable population.
| Biomarker |
Prevalence in Parent Population |
Source |
| PD-L1 TPS >= 50% (NSCLC) |
~30% |
KEYNOTE-024 screening data |
| EGFR mutations (NSCLC) |
~15% (Western), ~40% (Asian) |
TCGA, regional registries |
| HER2+ (breast cancer) |
~20% |
SEER, clinical databases |
| BRCA1/2 mutations (ovarian) |
~15-20% |
Population screening studies |
| MSI-H/dMMR (pan-tumor) |
~4-5% (all solid tumors) |
TCGA pan-cancer analysis |
| KRAS G12C (NSCLC) |
~13% |
AACR GENIE database |
Formula: Biomarker-positive = Eligible x Biomarker prevalence
Testing rate adjustment. Not all eligible patients are tested for the biomarker. Apply a testing rate multiplier: NGS adoption ~70-80% in academic centers, ~40-60% in community oncology (US, 2024). Testing rates are rising 5-10% annually in oncology.
Formula: Tested and positive = Eligible x Testing rate x Biomarker prevalence
Step 6 — Apply Line-of-Therapy Share
If the drug targets a specific treatment line, estimate the share of patients who reach that line.
| Line |
Typical Reach (Oncology) |
Typical Reach (Immunology) |
| 1L (first-line) |
100% (all treated patients) |
80-100% (mild may not be treated) |
| 2L |
40-60% |
50-70% |
| 3L |
20-35% |
30-50% |
| 4L+ |
10-15% |
15-25% |
Step 7 — Apply Geographic Weighting
Scale from primary geography to target markets.
| Market |
Share of Global Pharma Revenue |
Population Multiplier (from US base) |
| United States |
~45% of global |
1.0x |
| EU5 (DE, FR, UK, IT, ES) |
~20% of global |
0.8-1.0x (prevalence), 0.4-0.7x (revenue) |
| Japan |
~7% of global |
0.35-0.4x (prevalence), 0.3-0.5x (revenue) |
| China |
~8% of global (growing) |
3.0-4.0x (prevalence), 0.2-0.5x (revenue) |
| Rest of World |
~20% of global |
Variable |
Step 8 — Compile Population Waterfall
Output
PATIENT POPULATION ESTIMATE — [Indication]
Geography: [market]
Date: [assessment date]
POPULATION WATERFALL:
Total disease burden (prevalence/incidence): [N] Source: [GBD/SEER/etc]
x Diagnosis rate ([X]%): [N] Source: [citation]
x Treatment-seeking rate ([X]%): [N] Source: [citation]
x Treatment eligibility ([X]%): [N] Filters: [key exclusions]
x Biomarker prevalence ([X]%): [N] Biomarker: [name]
x Biomarker testing rate ([X]%): [N] Source: [citation]
x Line-of-therapy reach ([X]%): [N] Line: [1L/2L/3L+]
──────────────────────────────
TREATABLE ADDRESSABLE POPULATION: [N] [geography]
Global extrapolation (US x [multiplier]): [N]
CONFIDENCE RANGE:
Conservative (lower bound estimates): [N]
Base case: [N]
Optimistic (upper bound estimates): [N]
KEY ASSUMPTIONS:
1. [Most impactful assumption — which filter matters most]
2. [Second most impactful]
3. [Key trend that could change the estimate — e.g., biomarker testing adoption]
WORKED EXAMPLE BENCHMARK:
[Brief comparison to a marketed drug's actual patient volume in the same indication]
Worked example — Pembrolizumab 1L NSCLC (PD-L1 >= 50%, US):
- US NSCLC incidence: ~238,000/yr (ACS 2024)
- Metastatic at diagnosis: ~57% = ~136,000
- Diagnosis rate: ~90% = ~122,000
- Treatment-seeking (ECOG 0-1): ~70% = ~85,000
- PD-L1 TPS >= 50%: ~30% = ~25,500
- Testing rate: ~75% = ~19,100
- Treatable addressable population: ~19,000 patients/yr (US)
- Cross-check: Keytruda 1L NSCLC US revenue implies ~18,000-22,000 treated patients at ~$170K/yr — validates the funnel.
Error Handling
| Scenario |
Response |
| No reliable epidemiological data |
Use multiple lower-tier sources and triangulate; widen confidence range; flag data quality as key risk |
| Rare disease (<10,000 prevalence) |
Use rare disease registries (Orphanet, NORD); note that small populations amplify uncertainty at every funnel step; consider natural history study data |
| New biomarker without prevalence data |
Estimate from TCGA/GENIE genomic databases for oncology; for other TAs, use screening study data from clinical trials; flag as major assumption |
| Geographic data mismatch |
Clearly state which geography the base data represents; apply epidemiological adjustment factors for target geography; note ethnic/genetic prevalence differences (e.g., EGFR mutation rates 3x higher in Asia) |
| Rapidly evolving diagnostic landscape |
Date-stamp all testing rate assumptions; note directional trends; provide sensitivity analysis on testing rate |
Cross-Domain Connections
- Biotech-venture/peak-sales-forecaster: Primary consumer — patient population is the volume input to revenue modeling
- Biotech-venture/endpoint-selection: Biomarker prevalence affects trial feasibility and enrichment strategy
- Biotech-venture/clinical-development: Trial enrollment feasibility depends on population size at the eligible-patient step
- Biotech-venture/competitive-intelligence: Competitive entries affect line-of-therapy share estimates
1---2name: patient-population-sizer3description: Estimate addressable patient populations from epidemiological data, applying diagnostic rates, treatment-eligible filters, biomarker prevalence, and geographic adjustments to produce treatable population estimates for peak sales forecasting and clinical trial feasibility analysis.4---56# Patient Population Sizer — The Epidemiology-to-TAM Funnel78The gap between "50 million Americans have disease X" and "how many patients will take this drug" is where most revenue forecasts go wrong. A headline prevalence number is not a market. This skill builds a rigorous waterfall from total disease burden down to the treatable addressable population for a specific therapeutic asset, with sourced estimates at every step. The output feeds directly into peak-sales-forecaster as the patient volume input.910Every biotech pitch deck inflates patient numbers. This skill deflates them to reality.1112## How to Run1314### Input1516| Parameter | Required? | Example |17|---|---|---|18| Disease/indication | Yes | Non-small cell lung cancer (NSCLC) |19| Geography | Yes | US, EU5, Japan, Global |20| Line of therapy | Recommended | 1L metastatic, 2L+, adjuvant |21| Biomarker requirement | If applicable | PD-L1 TPS >= 50%, EGFR-mutant |22| Modality constraints | If applicable | IV infusion only (limits to infusion centers) |23| Competitive context | Recommended | Current SOC, anticipated entrants |2425### Steps2627> **Epidemiology data:** For sourced prevalence/incidence/survival (SEER 2026, GLOBOCAN 2022), cardiometabolic/CNS/immunology base rates (CDC, GBD), and a worked line-of-therapy attrition funnel, use `references/epidemiology-by-indication.md`. It is the single source of truth for biomarker/disease prevalence in the suite. Note global GBD figures understate US/EU addressable populations.2829#### Step 1 — Establish Disease Burden (Top of Funnel)3031Start with the broadest defensible population estimate. Use the data source hierarchy — always cite the highest-tier source available:3233| Tier | Source | Strengths | Limitations |34|---|---|---|---|35| 1 | GBD (Global Burden of Disease) | Standardized global methodology, 204 countries | Modeled estimates, 2-3yr lag |36| 2 | SEER / national cancer registries | Observed incidence, stage-specific | US/country-specific, cancer-only |37| 3 | Claims databases (Truven, Optum, IQVIA) | Real-world treatment patterns | Commercially insured bias, US-centric |38| 4 | Published epidemiological studies | Disease-specific depth | Single population, may be outdated |39| 5 | KOL estimates / company projections | Forward-looking | Unverifiable, potential bias |4041**Prevalence vs. incidence.** Use prevalence for chronic diseases where patients accumulate (e.g., rheumatoid arthritis, diabetes). Use incidence for diseases with rapid turnover (e.g., metastatic cancers where median survival < 2 years, acute infections). For oncology, annual incidence is typically the right starting point; for immunology, point prevalence.4243#### Step 2 — Apply Diagnostic Rate Filter4445Not all patients with a disease are diagnosed. This filter is often the largest single reduction in the funnel.4647| Disease Category | Typical Diagnosis Rate | Key Drivers |48|---|---|---|49| Common cancers (breast, lung, colon) | 85-95% | Screening programs, symptomatic presentation |50| Rare cancers | 70-85% | May be misdiagnosed; referral delays |51| Autoimmune diseases | 50-70% | Lengthy diagnostic odyssey, overlapping symptoms |52| Rare genetic diseases | 5-30% | Requires genetic testing, low awareness |53| Neurodegenerative (early-stage) | 30-60% | Gradual onset, normalization of symptoms |5455Formula: `Diagnosed population = Disease burden x Diagnosis rate`5657#### Step 3 — Apply Treatment-Seeking Filter5859Not all diagnosed patients seek or accept treatment.6061| Factor | Adjustment | Rationale |62|---|---|---|63| Asymptomatic early-stage disease | -20-40% | Watchful waiting, physician and patient reluctance |64| Stigmatized conditions | -15-30% | Mental health, substance abuse, sexual health |65| Elderly/frail patients | -10-20% | May not be offered aggressive therapy |66| Highly symptomatic, life-threatening | -5% or less | Strong motivation to treat |67| Disease with clear treatment guidelines | -5-10% | Guideline-concordant care drives treatment |6869Formula: `Treatment-seeking = Diagnosed x Treatment-seeking rate`7071#### Step 4 — Apply Treatment Eligibility Filters7273Clinical trial inclusion/exclusion criteria and label restrictions reduce the eligible population.7475| Filter | Typical Impact | Examples |76|---|---|---|77| Performance status (ECOG) | -10-20% | ECOG 0-1 required; ECOG 2+ excluded |78| Organ function requirements | -5-15% | Adequate hepatic/renal function |79| Prior therapy requirements | Variable | Must have failed 1L; may exclude pretreated |80| Contraindications | -5-10% | Autoimmune history for IO, cardiac for certain TKIs |81| Age restrictions | -2-5% | Pediatric exclusion, upper age limits |82| Comorbidities | -10-20% | Uncontrolled diabetes, active infections |8384Formula: `Eligible = Treatment-seeking x (1 - sum of exclusion rates)`8586#### Step 5 — Apply Biomarker Prevalence Filter8788If the drug requires biomarker selection, this can dramatically reduce the addressable population.8990| Biomarker | Prevalence in Parent Population | Source |91|---|---|---|92| PD-L1 TPS >= 50% (NSCLC) | ~30% | KEYNOTE-024 screening data |93| EGFR mutations (NSCLC) | ~15% (Western), ~40% (Asian) | TCGA, regional registries |94| HER2+ (breast cancer) | ~20% | SEER, clinical databases |95| BRCA1/2 mutations (ovarian) | ~15-20% | Population screening studies |96| MSI-H/dMMR (pan-tumor) | ~4-5% (all solid tumors) | TCGA pan-cancer analysis |97| KRAS G12C (NSCLC) | ~13% | AACR GENIE database |9899Formula: `Biomarker-positive = Eligible x Biomarker prevalence`100101**Testing rate adjustment.** Not all eligible patients are tested for the biomarker. Apply a testing rate multiplier: NGS adoption ~70-80% in academic centers, ~40-60% in community oncology (US, 2024). Testing rates are rising 5-10% annually in oncology.102103Formula: `Tested and positive = Eligible x Testing rate x Biomarker prevalence`104105#### Step 6 — Apply Line-of-Therapy Share106107If the drug targets a specific treatment line, estimate the share of patients who reach that line.108109| Line | Typical Reach (Oncology) | Typical Reach (Immunology) |110|---|---|---|111| 1L (first-line) | 100% (all treated patients) | 80-100% (mild may not be treated) |112| 2L | 40-60% | 50-70% |113| 3L | 20-35% | 30-50% |114| 4L+ | 10-15% | 15-25% |115116#### Step 7 — Apply Geographic Weighting117118Scale from primary geography to target markets.119120| Market | Share of Global Pharma Revenue | Population Multiplier (from US base) |121|---|---|---|122| United States | ~45% of global | 1.0x |123| EU5 (DE, FR, UK, IT, ES) | ~20% of global | 0.8-1.0x (prevalence), 0.4-0.7x (revenue) |124| Japan | ~7% of global | 0.35-0.4x (prevalence), 0.3-0.5x (revenue) |125| China | ~8% of global (growing) | 3.0-4.0x (prevalence), 0.2-0.5x (revenue) |126| Rest of World | ~20% of global | Variable |127128#### Step 8 — Compile Population Waterfall129130### Output131132```133PATIENT POPULATION ESTIMATE — [Indication]134Geography: [market]135Date: [assessment date]136137POPULATION WATERFALL:138 Total disease burden (prevalence/incidence): [N] Source: [GBD/SEER/etc]139 x Diagnosis rate ([X]%): [N] Source: [citation]140 x Treatment-seeking rate ([X]%): [N] Source: [citation]141 x Treatment eligibility ([X]%): [N] Filters: [key exclusions]142 x Biomarker prevalence ([X]%): [N] Biomarker: [name]143 x Biomarker testing rate ([X]%): [N] Source: [citation]144 x Line-of-therapy reach ([X]%): [N] Line: [1L/2L/3L+]145 ──────────────────────────────146 TREATABLE ADDRESSABLE POPULATION: [N] [geography]147148 Global extrapolation (US x [multiplier]): [N]149150CONFIDENCE RANGE:151 Conservative (lower bound estimates): [N]152 Base case: [N]153 Optimistic (upper bound estimates): [N]154155KEY ASSUMPTIONS:156 1. [Most impactful assumption — which filter matters most]157 2. [Second most impactful]158 3. [Key trend that could change the estimate — e.g., biomarker testing adoption]159160WORKED EXAMPLE BENCHMARK:161 [Brief comparison to a marketed drug's actual patient volume in the same indication]162```163164**Worked example — Pembrolizumab 1L NSCLC (PD-L1 >= 50%, US):**165- US NSCLC incidence: ~238,000/yr (ACS 2024)166- Metastatic at diagnosis: ~57% = ~136,000167- Diagnosis rate: ~90% = ~122,000168- Treatment-seeking (ECOG 0-1): ~70% = ~85,000169- PD-L1 TPS >= 50%: ~30% = ~25,500170- Testing rate: ~75% = ~19,100171- Treatable addressable population: ~19,000 patients/yr (US)172- Cross-check: Keytruda 1L NSCLC US revenue implies ~18,000-22,000 treated patients at ~$170K/yr — validates the funnel.173174### Error Handling175176| Scenario | Response |177|---|---|178| No reliable epidemiological data | Use multiple lower-tier sources and triangulate; widen confidence range; flag data quality as key risk |179| Rare disease (<10,000 prevalence) | Use rare disease registries (Orphanet, NORD); note that small populations amplify uncertainty at every funnel step; consider natural history study data |180| New biomarker without prevalence data | Estimate from TCGA/GENIE genomic databases for oncology; for other TAs, use screening study data from clinical trials; flag as major assumption |181| Geographic data mismatch | Clearly state which geography the base data represents; apply epidemiological adjustment factors for target geography; note ethnic/genetic prevalence differences (e.g., EGFR mutation rates 3x higher in Asia) |182| Rapidly evolving diagnostic landscape | Date-stamp all testing rate assumptions; note directional trends; provide sensitivity analysis on testing rate |183184## Cross-Domain Connections185186- **Biotech-venture/peak-sales-forecaster**: Primary consumer — patient population is the volume input to revenue modeling187- **Biotech-venture/endpoint-selection**: Biomarker prevalence affects trial feasibility and enrichment strategy188- **Biotech-venture/clinical-development**: Trial enrollment feasibility depends on population size at the eligible-patient step189- **Biotech-venture/competitive-intelligence**: Competitive entries affect line-of-therapy share estimates