1---2name: assumptions-and-limitations3description: Every statistical method has assumptions. When assumptions are violated, results may be biased, inefficient, or misleading.4---5# Assumptions and Limitations67## Overview89Every statistical method has assumptions. When assumptions are violated, results may be biased, inefficient, or misleading. This reference lists key assumptions by method, how to check them, and what to do when they fail.1011---1213## Linear Regression (OLS)1415### Assumptions1617| Assumption | What It Means | How to Check | Consequence if Violated |18|------------|---------------|--------------|------------------------|19| Linearity | Relationship is linear | Residual vs fitted plot | Biased estimates |20| Independence | Errors are independent | Study design, Durbin-Watson | Wrong SEs, invalid inference |21| Homoscedasticity | Constant error variance | Residual plot, Breusch-Pagan | Inefficient estimates, wrong SEs |22| Normality | Residuals are normal | Q-Q plot, Shapiro-Wilk | Invalid inference (small n) |23| No multicollinearity | Predictors not highly correlated | VIF (>10 concerning) | Unstable estimates |2425### Remedies2627| Violation | Options |28|-----------|---------|29| Non-linearity | Transform variables, add polynomial terms, use GAM |30| Heteroscedasticity | Robust (HC) standard errors, WLS, transform outcome |31| Non-independence | Cluster-robust SEs, mixed models, GEE |32| Non-normality | Bootstrap CIs, transform outcome (large n usually OK) |33| Multicollinearity | Remove redundant predictors, combine, regularization |3435---3637## Logistic Regression3839### Assumptions4041| Assumption | What It Means | How to Check | Consequence if Violated |42|------------|---------------|--------------|------------------------|43| Linearity in log-odds | Linear relationship on logit scale | Box-Tidwell test | Biased estimates |44| Independence | Observations are independent | Study design | Wrong SEs |45| No perfect separation | Both outcomes exist for all predictor values | Check for infinite coefficients | Model fails to converge |46| No multicollinearity | Predictors not highly correlated | VIF | Unstable estimates |4748### Remedies4950| Violation | Options |51|-----------|---------|52| Non-linearity | Splines, categorize continuous variables |53| Non-independence | GEE, mixed logistic models, cluster-robust SEs |54| Perfect separation | Firth's penalized logistic regression, exact logistic |55| Multicollinearity | Remove predictors, regularization (LASSO) |5657### Additional Considerations5859- **Rare events:** When outcome prevalence <10%, consider exact logistic or Firth's method60- **Sample size rule of thumb:** ~10-20 events per predictor variable6162---6364## Cox Proportional Hazards6566### Assumptions6768| Assumption | What It Means | How to Check | Consequence if Violated |69|------------|---------------|--------------|------------------------|70| Proportional hazards | Hazard ratio constant over time | Schoenfeld residuals, log-log plot | HR is time-averaged, misleading |71| Non-informative censoring | Censoring unrelated to outcome | Cannot test; assess plausibility | Biased estimates |72| Linearity | Continuous covariates linear in log-hazard | Martingale residuals | Biased estimates |73| Independence | Observations independent | Study design | Wrong SEs |7475### Checking Proportional Hazards76771. **Schoenfeld residual test:** `cox.zph()` in R, `proportional_hazard_test` in lifelines782. **Log-log plot:** Parallel lines indicate PH holds793. **Include time interaction:** If significant, PH violated8081### Remedies8283| Violation | Options |84|-----------|---------|85| PH violated (categorical) | Stratified Cox |86| PH violated (continuous) | Time-varying coefficient, piecewise Cox |87| PH generally problematic | AFT models, RMST |88| Informative censoring | Sensitivity analysis, consider competing risks |89| Non-linearity | Splines, categorize |9091---9293## Poisson Regression9495### Assumptions9697| Assumption | What It Means | How to Check | Consequence if Violated |98|------------|---------------|--------------|------------------------|99| Mean = Variance | Equidispersion | Compare mean and variance, deviance/df | Wrong SEs (usually too small) |100| Independence | Events are independent | Study design | Wrong SEs |101| Log-linear relationship | Correct functional form | Residual plots | Biased estimates |102103### Checking Overdispersion104105- Dispersion ratio = Pearson χ² / df (or deviance / df)106- Ratio > 1.5 suggests overdispersion107108### Remedies109110| Violation | Options |111|-----------|---------|112| Overdispersion | Negative binomial, quasi-Poisson |113| Excess zeros | Zero-inflated Poisson, hurdle model |114| Non-independence | GEE, mixed Poisson models |115116---117118## IPTW / Propensity Score Methods119120### Assumptions121122| Assumption | What It Means | How to Check | Consequence if Violated |123|------------|---------------|--------------|------------------------|124| No unmeasured confounding | All confounders in PS model | Cannot test; defend with theory | Biased causal estimate |125| Positivity | All covariate patterns have treated & untreated | Check PS distribution overlap | Extreme weights, instability |126| Correct PS model | Propensity model well-specified | Covariate balance after weighting | Residual confounding |127| Consistency | Well-defined treatment | Conceptual | Effect not interpretable |128129### Checking Balance130131After weighting or matching, check:132- Standardized mean differences (SMD < 0.1 ideal)133- Variance ratios (0.5-2.0 acceptable)134- Visual: distribution overlap plots135136### Remedies137138| Violation | Options |139|-----------|---------|140| Poor overlap | Trim extreme weights, match instead of weight |141| Residual imbalance | Adjust PS model, add interactions |142| Extreme weights | Truncate or stabilize weights |143| Unmeasured confounding | Sensitivity analysis (E-value), different design |144145---146147## Mixed Effects Models148149### Assumptions150151| Assumption | What It Means | How to Check | Consequence if Violated |152|------------|---------------|--------------|------------------------|153| Random effects normal | Random effects follow normal distribution | Q-Q plot of BLUPs | Usually robust |154| Residuals normal & homoscedastic | Standard regression assumptions | Residual plots | Biased SEs |155| Correct random structure | Random effects specified correctly | Model comparison (AIC/BIC) | Biased estimates or SEs |156| Independence of clusters | Clusters are independent | Study design | Wrong inference |157158### Remedies159160| Violation | Options |161|-----------|---------|162| Non-normal random effects | Bootstrap, robust SEs |163| Heteroscedasticity | Model variance by group |164| Uncertain random structure | Compare models, keep parsimonious |165166---167168## Machine Learning Models169170### Considerations (Not Traditional Assumptions)171172| Issue | What It Means | How to Address |173|-------|---------------|----------------|174| Overfitting | Model memorizes training data | Cross-validation, regularization, early stopping |175| Data leakage | Future info in training | Careful feature engineering, temporal splits |176| Class imbalance | Rare outcome | SMOTE, class weights, AUPRC instead of AUROC |177| Feature scaling | Algorithms sensitive to scale | Standardize/normalize (tree-based models don't need) |178| Missing data | Incomplete features | Imputation, models that handle missingness |179180### Validation Requirements181182- **Always** use held-out test set or cross-validation183- **Never** report training metrics as final performance184- **Patient-level splits** to avoid data leakage across admissions185186---187188## Universal Considerations189190### Sample Size191192| Method | Rule of Thumb |193|--------|---------------|194| Linear regression | 10-20 observations per predictor |195| Logistic regression | 10-20 events per predictor |196| Cox regression | 10-20 events per predictor |197| ML models | Hundreds to thousands; more for complex models |198199### Missing Data200201| Mechanism | Description | Appropriate Approach |202|-----------|-------------|---------------------|203| MCAR | Missingness completely random | Complete case (but loses power) |204| MAR | Missingness depends on observed data | Multiple imputation |205| MNAR | Missingness depends on missing value | Sensitivity analysis, pattern mixture |206207### Multiple Testing208209If testing multiple hypotheses:210- Pre-specify primary outcome211- Bonferroni: α / number of tests (conservative)212- Benjamini-Hochberg: Controls false discovery rate (less conservative)213- Report all tests performed, not just significant ones214215---216217## Quick Reference: Assumption → Check → Fix218219| Assumption | Quick Check | Quick Fix |220|------------|-------------|-----------|221| Linearity | Residual vs fitted plot | Transform, splines, GAM |222| Homoscedasticity | Residual spread pattern | Robust SEs, WLS |223| Normality | Q-Q plot | Bootstrap (or ignore if n>30) |224| Independence | Study design | Cluster methods, GEE |225| PH (Cox) | Schoenfeld test | Stratify, time-varying, AFT |226| Overdispersion | Mean vs variance | Negative binomial |227| PS overlap | Weight distribution | Trim, match, bound |228| Multicollinearity | VIF | Remove, combine, regularize |