Question Taxonomy
Overview
The type of research question determines the appropriate analytical approach. Misclassifying a question leads to methods that don't answer what you actually want to know.
Primary Question Types
1. Prediction
"Can we accurately forecast outcome Y given information X?"
- Goal: Maximize predictive accuracy for new, unseen cases
- Focus: Model performance (discrimination, calibration)
- Coefficients: Not the primary interest; may be uninterpretable
- Causation: Not claimed or required
Examples:
- "Can we predict ICU mortality at admission?"
- "Which patients will develop AKI in the next 24 hours?"
- "What is the risk of readmission for this patient?"
Appropriate methods: Logistic regression, random forest, gradient boosting, neural networks, penalized regression
Key considerations:
- Features must be available at prediction time (no future information leakage)
- Validation on truly held-out data is essential
- Calibration matters for clinical use, not just discrimination
2. Inference (Association)
"Is X associated with Y, accounting for confounders?"
- Goal: Estimate the relationship between exposure and outcome
- Focus: Effect size, confidence interval, statistical significance
- Coefficients: Primary interest; must be interpretable
- Causation: Claimed only with strong design (RCT, quasi-experimental)
Examples:
- "Is vasopressor choice associated with mortality?"
- "Do patients with diabetes have longer ICU stays?"
- "Is early mobilization associated with reduced delirium?"
Appropriate methods: Linear/logistic regression, Cox regression, GEE, mixed models
Key considerations:
- Confounder identification and adjustment
- Model specification (correct functional form)
- Effect modification / interaction
3. Causal Inference
"Does X cause Y?"
- Goal: Estimate the causal effect of an intervention
- Focus: What would happen if we changed X?
- Requires: Strong assumptions (exchangeability, positivity, consistency)
- Gold standard: Randomized controlled trial
Examples:
- "Does early intubation reduce mortality compared to delayed intubation?"
- "What is the effect of a restrictive transfusion strategy on outcomes?"
Appropriate methods: IPTW, propensity score matching, instrumental variables, difference-in-differences, regression discontinuity
Key considerations:
- Causal assumptions must be explicitly stated and defended
- Unmeasured confounding is always a threat in observational data
- Sensitivity analyses for violations are essential
4. Description
"What is the distribution/frequency of X?"
- Goal: Characterize a population or phenomenon
- Focus: Summary statistics, patterns, trends
- No exposure-outcome relationship tested
Examples:
- "What is the 30-day mortality rate in sepsis patients?"
- "How has ICU length of stay changed over time?"
- "What are the most common diagnoses in our cohort?"
Appropriate methods: Descriptive statistics, data visualization, trend analysis
Key considerations:
- Clearly define the population
- Report variability (SD, IQR, range), not just central tendency
- Consider selection bias in who enters the cohort
5. Clustering / Subgroup Discovery
"Are there natural subgroups within this population?"
- Goal: Identify latent structure or phenotypes
- Focus: Group membership, cluster characteristics
- Unsupervised: No predefined outcome
Examples:
- "Are there distinct sepsis phenotypes?"
- "Can we identify patient subgroups with different treatment responses?"
Appropriate methods: K-means, hierarchical clustering, latent class analysis, Gaussian mixture models
Key considerations:
- Cluster validity (internal and external)
- Clinical interpretability of clusters
- Stability across samples
Secondary Distinctions
Exploratory vs. Confirmatory
| Aspect |
Exploratory |
Confirmatory |
| Hypothesis |
Generated from data |
Pre-specified |
| Multiple testing |
Expected |
Must be controlled |
| Reporting |
All findings |
Primary outcome focus |
| Replication |
Needed |
Is the replication |
Cross-sectional vs. Longitudinal
| Aspect |
Cross-sectional |
Longitudinal |
| Timepoints |
Single |
Multiple |
| Causation |
Cannot establish temporal order |
Can establish temporal sequence |
| Analysis |
Standard regression |
Mixed models, GEE, survival |
Time-to-Event vs. Binary/Continuous
| Outcome Type |
Characteristics |
Methods |
| Binary |
Yes/no at fixed time |
Logistic regression, chi-square |
| Continuous |
Measured value |
Linear regression, t-test, ANOVA |
| Time-to-event |
When did it happen? + censoring |
Cox, Kaplan-Meier, competing risks |
| Count |
Number of events |
Poisson, negative binomial |
Decision Aid: What Type is My Question?
START: What do you want to learn?
│
├─ "How well can we forecast Y?"
│ └─ → PREDICTION
│
├─ "Is X related to Y?"
│ ├─ "...and I want to claim X causes Y"
│ │ └─ → CAUSAL INFERENCE (requires strong assumptions)
│ └─ "...adjusting for confounders"
│ └─ → INFERENCE (association)
│
├─ "What does Y look like in this population?"
│ └─ → DESCRIPTION
│
└─ "Are there natural groups in my data?"
└─ → CLUSTERING
Common Misclassifications
Prediction disguised as inference
- Red flag: "We found that feature X has high importance in predicting Y, therefore X is associated with Y"
- Problem: Variable importance ≠ causal effect; prediction models optimize for accuracy, not interpretability
Association claimed as causation
- Red flag: "Patients who received treatment X had lower mortality, so X reduces mortality"
- Problem: Without randomization or causal methods, confounding by indication is likely
Description without context
- Red flag: "Mortality was 25% in our cohort"
- Problem: Without comparison group or benchmark, the number is uninterpretable
Implications for Analysis Planning
| Question Type |
What to prioritize |
| Prediction |
Validation strategy, calibration, feature engineering |
| Inference |
Confounder identification, model specification, CIs |
| Causal |
Assumptions, sensitivity analysis, design elements |
| Description |
Population definition, complete reporting |
| Clustering |
Validity metrics, interpretability, stability |
1---2name: question-taxonomy3description: The type of research question determines the appropriate analytical approach. Misclassifying a question leads to methods that don't answer what you actually want to know.4---5# Question Taxonomy67## Overview89The type of research question determines the appropriate analytical approach. Misclassifying a question leads to methods that don't answer what you actually want to know.1011---1213## Primary Question Types1415### 1. Prediction1617**"Can we accurately forecast outcome Y given information X?"**1819- Goal: Maximize predictive accuracy for new, unseen cases20- Focus: Model performance (discrimination, calibration)21- Coefficients: Not the primary interest; may be uninterpretable22- Causation: Not claimed or required2324**Examples:**25- "Can we predict ICU mortality at admission?"26- "Which patients will develop AKI in the next 24 hours?"27- "What is the risk of readmission for this patient?"2829**Appropriate methods:** Logistic regression, random forest, gradient boosting, neural networks, penalized regression3031**Key considerations:**32- Features must be available at prediction time (no future information leakage)33- Validation on truly held-out data is essential34- Calibration matters for clinical use, not just discrimination3536---3738### 2. Inference (Association)3940**"Is X associated with Y, accounting for confounders?"**4142- Goal: Estimate the relationship between exposure and outcome43- Focus: Effect size, confidence interval, statistical significance44- Coefficients: Primary interest; must be interpretable45- Causation: Claimed only with strong design (RCT, quasi-experimental)4647**Examples:**48- "Is vasopressor choice associated with mortality?"49- "Do patients with diabetes have longer ICU stays?"50- "Is early mobilization associated with reduced delirium?"5152**Appropriate methods:** Linear/logistic regression, Cox regression, GEE, mixed models5354**Key considerations:**55- Confounder identification and adjustment56- Model specification (correct functional form)57- Effect modification / interaction5859---6061### 3. Causal Inference6263**"Does X cause Y?"**6465- Goal: Estimate the causal effect of an intervention66- Focus: What would happen if we changed X?67- Requires: Strong assumptions (exchangeability, positivity, consistency)68- Gold standard: Randomized controlled trial6970**Examples:**71- "Does early intubation reduce mortality compared to delayed intubation?"72- "What is the effect of a restrictive transfusion strategy on outcomes?"7374**Appropriate methods:** IPTW, propensity score matching, instrumental variables, difference-in-differences, regression discontinuity7576**Key considerations:**77- Causal assumptions must be explicitly stated and defended78- Unmeasured confounding is always a threat in observational data79- Sensitivity analyses for violations are essential8081---8283### 4. Description8485**"What is the distribution/frequency of X?"**8687- Goal: Characterize a population or phenomenon88- Focus: Summary statistics, patterns, trends89- No exposure-outcome relationship tested9091**Examples:**92- "What is the 30-day mortality rate in sepsis patients?"93- "How has ICU length of stay changed over time?"94- "What are the most common diagnoses in our cohort?"9596**Appropriate methods:** Descriptive statistics, data visualization, trend analysis9798**Key considerations:**99- Clearly define the population100- Report variability (SD, IQR, range), not just central tendency101- Consider selection bias in who enters the cohort102103---104105### 5. Clustering / Subgroup Discovery106107**"Are there natural subgroups within this population?"**108109- Goal: Identify latent structure or phenotypes110- Focus: Group membership, cluster characteristics111- Unsupervised: No predefined outcome112113**Examples:**114- "Are there distinct sepsis phenotypes?"115- "Can we identify patient subgroups with different treatment responses?"116117**Appropriate methods:** K-means, hierarchical clustering, latent class analysis, Gaussian mixture models118119**Key considerations:**120- Cluster validity (internal and external)121- Clinical interpretability of clusters122- Stability across samples123124---125126## Secondary Distinctions127128### Exploratory vs. Confirmatory129130| Aspect | Exploratory | Confirmatory |131|--------|-------------|--------------|132| Hypothesis | Generated from data | Pre-specified |133| Multiple testing | Expected | Must be controlled |134| Reporting | All findings | Primary outcome focus |135| Replication | Needed | Is the replication |136137### Cross-sectional vs. Longitudinal138139| Aspect | Cross-sectional | Longitudinal |140|--------|-----------------|--------------|141| Timepoints | Single | Multiple |142| Causation | Cannot establish temporal order | Can establish temporal sequence |143| Analysis | Standard regression | Mixed models, GEE, survival |144145### Time-to-Event vs. Binary/Continuous146147| Outcome Type | Characteristics | Methods |148|--------------|-----------------|---------|149| Binary | Yes/no at fixed time | Logistic regression, chi-square |150| Continuous | Measured value | Linear regression, t-test, ANOVA |151| Time-to-event | When did it happen? + censoring | Cox, Kaplan-Meier, competing risks |152| Count | Number of events | Poisson, negative binomial |153154---155156## Decision Aid: What Type is My Question?157158```159START: What do you want to learn?160│161├─ "How well can we forecast Y?"162│ └─ → PREDICTION163│164├─ "Is X related to Y?"165│ ├─ "...and I want to claim X causes Y"166│ │ └─ → CAUSAL INFERENCE (requires strong assumptions)167│ └─ "...adjusting for confounders"168│ └─ → INFERENCE (association)169│170├─ "What does Y look like in this population?"171│ └─ → DESCRIPTION172│173└─ "Are there natural groups in my data?"174 └─ → CLUSTERING175```176177---178179## Common Misclassifications180181### Prediction disguised as inference182- **Red flag:** "We found that feature X has high importance in predicting Y, therefore X is associated with Y"183- **Problem:** Variable importance ≠ causal effect; prediction models optimize for accuracy, not interpretability184185### Association claimed as causation186- **Red flag:** "Patients who received treatment X had lower mortality, so X reduces mortality"187- **Problem:** Without randomization or causal methods, confounding by indication is likely188189### Description without context190- **Red flag:** "Mortality was 25% in our cohort"191- **Problem:** Without comparison group or benchmark, the number is uninterpretable192193---194195## Implications for Analysis Planning196197| Question Type | What to prioritize |198|---------------|-------------------|199| Prediction | Validation strategy, calibration, feature engineering |200| Inference | Confounder identification, model specification, CIs |201| Causal | Assumptions, sensitivity analysis, design elements |202| Description | Population definition, complete reporting |203| Clustering | Validity metrics, interpretability, stability |