Bambi Python
Translate a statistical estimand into an explicit formula, likelihood family,
link, contrast/group structure, and prior contract; inspect the built model;
then validate inference and predictions rather than treating .fit() as done.
Boundary
Use Bambi when formula-driven Bayesian regression or multilevel modeling is the
primary job. Use PyMC for a custom probability graph that cannot be expressed
cleanly through Bambi, and ArviZ for analysis that starts from an existing
InferenceData. Bambi uses a PyMC backend, but backend implementation details
must not silently replace the formula-level contract.
Know the objects
| Object |
Meaning |
Use it for |
| Formula |
Response plus common and group-specific linear predictor terms; optional additional formulas target other distribution parameters. |
Declaring estimands and hierarchical structure. |
| Term |
A transformed predictor, interaction, intercept, or group-specific coefficient produced by formula parsing. |
Understanding the actual design matrix and coefficient interpretation. |
Family |
Likelihood plus link(s) connecting distribution parameters to predictors. |
Matching outcome support and data-generating process. |
Prior |
Distribution specification for a term or auxiliary parameter, optionally nested/hierarchical. |
Encoding domain scale and regularization. |
Model |
Formula, data, family, priors, design matrices, settings, and compiled PyMC backend. |
Building, fitting, and predicting one coherent regression. |
| Backend model |
The generated PyMC model after build(). |
Low-level inspection and supported sampler/backend configuration. |
InferenceData |
Labeled posterior, sample statistics, log likelihood, observed data, and predictive groups returned/extended by fitting and prediction. |
Diagnostics, persistence, and ArviZ handoff. |
Common effects estimate population-level coefficients. Group-specific terms such
as (1 + x | group) estimate varying intercepts/slopes with partial pooling; the
vertical bar is not a bitwise operation or ordinary interaction. Read the
formula and model contract before changing term
syntax, coding, or hierarchy.
Ordered workflow
- State the outcome support, observation unit, estimand, likelihood, link,
exposure/trials, repeated/group structure, missing-data policy, and prediction
population.
- Validate the pandas data before model construction: response support,
category levels/order, group IDs, numeric units, missing rows, and one row per
declared observation unit.
- Write the smallest formula that represents the estimand. Add interactions,
splines/transforms, offsets, and group-specific terms only under explicit
scientific meaning.
- Choose
Family and link from the sampling process, not from the outcome's
Python dtype. Include auxiliary/distributional formulas when dispersion or
another parameter genuinely varies.
- Inspect the printed/built model, design terms, family/link, observations
retained, categorical reference levels, priors, and parameter names.
- Replace consequential automatic priors with explicit domain-scale priors.
Run prior predictive simulation before fitting.
- Fit with recorded backend, chains, warmup, draws, seed, and sampler settings.
Diagnose the returned
InferenceData using PyMC/ArviZ evidence.
- Run posterior predictive checks. For new data, preserve formula columns,
factor levels, transforms, and group policy; store predictions separately.
Decision map
| Condition |
Action |
| Continuous approximately symmetric outcome |
Gaussian may be a candidate; verify tails, bounds, and variance behavior. |
| Binary outcome |
Bernoulli family with a justified link; validate values and event unit. |
| Counts |
Poisson only if mean/variance and zero behavior fit; otherwise test overdispersion/zero structure with a supported family/model. |
| Binomial successes out of trials |
Supply the binomial trial contract; do not model counts as unconstrained Poisson by habit. |
| Repeated observations by group |
Add group-specific intercept/slopes required by dependence and estimand. |
| Few observations per group or funnel diagnostics |
Prefer/test non-centered group effects; do not eliminate pooling. |
| Prior scale has domain meaning |
Set explicit Prior; do not rely on automatic scaling invisibly. |
| New data contains unseen group levels |
Apply an explicit supported new-group policy and uncertainty interpretation; never silently map to an existing group. |
| Custom latent graph exceeds formula/family system |
Drop to a hand-built PyMC model and use that skill. |
Read families, priors, and inference before
accepting defaults or changing the backend.
Canonical hierarchical anchor
import arviz as az
import bambi as bmb
import pandas as pd
def fit_reaction(data: pd.DataFrame, seed: int):
required = {"reaction", "days", "subject"}
if missing := required.difference(data.columns):
raise ValueError(f"missing columns: {sorted(missing)}")
if data[list(required)].isna().any().any():
raise ValueError("missing model values require an explicit policy")
priors = {
"days": bmb.Prior("Normal", mu=0, sigma=5),
"1|subject": bmb.Prior("Normal", mu=0, sigma=bmb.Prior("HalfNormal", sigma=10)),
}
model = bmb.Model(
"reaction ~ 1 + days + (1 + days | subject)",
data,
family="gaussian",
priors=priors,
dropna=False,
)
model.build()
prior = model.prior_predictive(random_seed=seed)
idata = model.fit(chains=4, random_seed=seed)
summary = az.summary(idata)
return model, prior, idata, summary
Prior names and signatures drift; inspect the built model in the installed
Bambi version. The chosen numeric scales are examples, not domain defaults.
Non-negotiable rules
- Do not let Bambi drop rows silently. Reject missing model columns or opt into
listwise deletion and report the exact retained observation set.
- Formula syntax does not validate causal identification. Confounding,
post-treatment variables, selection, and measurement error require a model
justified outside the formula parser.
- Inspect intercept inclusion, categorical reference levels, interactions, and
group terms. Coefficient meaning changes with coding and predictor centering.
- Automatic prior scaling is a convenience, not a domain prior. Record whether
auto_scale and predictor centering are active and inspect resulting priors.
- A group-specific term requires a meaningful population of groups. Do not use
fixed dummy effects merely to avoid partial pooling, and do not add random
slopes unsupported by data without diagnosing geometry.
.fit() returning is not acceptance. Inspect divergences, R-hat, ESS, MCSE,
traces, prior/posterior predictive checks, and parameterization.
- Predictions on new data must reuse training transforms and category coding.
State whether they include observation noise and group-specific effects.
- Model comparison requires compatible observations and pointwise log
likelihood; use ArviZ and inspect Pareto diagnostics.
Use testing and version grounding to inspect
the installed formula, fit, predict, and backend contracts.
Completion gate
Do not declare completion until the formula matches the estimand; response
family/link and auxiliary parameters match support; retained rows and category/
group encodings are explicit; built terms and priors are inspected; prior
predictive and simulated recovery checks pass; inference has no unresolved
critical diagnostics; posterior predictive checks target domain behavior; new-
data prediction semantics and unseen groups are tested; InferenceData groups
and coordinates are correct; and skipped checks or version uncertainty are
reported.
References
- Formula and model contract
- Families, priors, and inference
- Testing and version grounding
1---2name: bambi-python3description: Use for writing, reviewing, debugging, testing, or diagnosing Bayesian regression and hierarchical models built with Bambi formulas, Model, Family/Likelihood/Link, Prior, fit, prior predictive, and predict. Trigger on common versus group-specific terms, categorical coding, family/link choice, automatic prior scaling, missing rows, PyMC backend settings, and InferenceData predictions. Do not use for hand-built PyMC graphs, NumPyro programs, ArviZ-only analysis of existing draws, frequentist statsmodels formulas, or generic pandas work.4---56# Bambi Python78Translate a statistical estimand into an explicit formula, likelihood family,9link, contrast/group structure, and prior contract; inspect the built model;10then validate inference and predictions rather than treating `.fit()` as done.1112## Boundary1314Use Bambi when formula-driven Bayesian regression or multilevel modeling is the15primary job. Use PyMC for a custom probability graph that cannot be expressed16cleanly through Bambi, and ArviZ for analysis that starts from an existing17`InferenceData`. Bambi uses a PyMC backend, but backend implementation details18must not silently replace the formula-level contract.1920## Know the objects2122| Object | Meaning | Use it for |23|---|---|---|24| Formula | Response plus common and group-specific linear predictor terms; optional additional formulas target other distribution parameters. | Declaring estimands and hierarchical structure. |25| Term | A transformed predictor, interaction, intercept, or group-specific coefficient produced by formula parsing. | Understanding the actual design matrix and coefficient interpretation. |26| `Family` | Likelihood plus link(s) connecting distribution parameters to predictors. | Matching outcome support and data-generating process. |27| `Prior` | Distribution specification for a term or auxiliary parameter, optionally nested/hierarchical. | Encoding domain scale and regularization. |28| `Model` | Formula, data, family, priors, design matrices, settings, and compiled PyMC backend. | Building, fitting, and predicting one coherent regression. |29| Backend model | The generated PyMC model after `build()`. | Low-level inspection and supported sampler/backend configuration. |30| `InferenceData` | Labeled posterior, sample statistics, log likelihood, observed data, and predictive groups returned/extended by fitting and prediction. | Diagnostics, persistence, and ArviZ handoff. |3132Common effects estimate population-level coefficients. Group-specific terms such33as `(1 + x | group)` estimate varying intercepts/slopes with partial pooling; the34vertical bar is not a bitwise operation or ordinary interaction. Read [the35formula and model contract](references/object-model.md) before changing term36syntax, coding, or hierarchy.3738## Ordered workflow39401. State the outcome support, observation unit, estimand, likelihood, link,41 exposure/trials, repeated/group structure, missing-data policy, and prediction42 population.432. Validate the pandas data before model construction: response support,44 category levels/order, group IDs, numeric units, missing rows, and one row per45 declared observation unit.463. Write the smallest formula that represents the estimand. Add interactions,47 splines/transforms, offsets, and group-specific terms only under explicit48 scientific meaning.494. Choose `Family` and link from the sampling process, not from the outcome's50 Python dtype. Include auxiliary/distributional formulas when dispersion or51 another parameter genuinely varies.525. Inspect the printed/built model, design terms, family/link, observations53 retained, categorical reference levels, priors, and parameter names.546. Replace consequential automatic priors with explicit domain-scale priors.55 Run prior predictive simulation before fitting.567. Fit with recorded backend, chains, warmup, draws, seed, and sampler settings.57 Diagnose the returned `InferenceData` using PyMC/ArviZ evidence.588. Run posterior predictive checks. For new data, preserve formula columns,59 factor levels, transforms, and group policy; store predictions separately.6061## Decision map6263| Condition | Action |64|---|---|65| Continuous approximately symmetric outcome | Gaussian may be a candidate; verify tails, bounds, and variance behavior. |66| Binary outcome | Bernoulli family with a justified link; validate values and event unit. |67| Counts | Poisson only if mean/variance and zero behavior fit; otherwise test overdispersion/zero structure with a supported family/model. |68| Binomial successes out of trials | Supply the binomial trial contract; do not model counts as unconstrained Poisson by habit. |69| Repeated observations by group | Add group-specific intercept/slopes required by dependence and estimand. |70| Few observations per group or funnel diagnostics | Prefer/test non-centered group effects; do not eliminate pooling. |71| Prior scale has domain meaning | Set explicit `Prior`; do not rely on automatic scaling invisibly. |72| New data contains unseen group levels | Apply an explicit supported new-group policy and uncertainty interpretation; never silently map to an existing group. |73| Custom latent graph exceeds formula/family system | Drop to a hand-built PyMC model and use that skill. |7475Read [families, priors, and inference](references/inference-priors.md) before76accepting defaults or changing the backend.7778## Canonical hierarchical anchor7980```python81import arviz as az82import bambi as bmb83import pandas as pd848586def fit_reaction(data: pd.DataFrame, seed: int):87 required = {"reaction", "days", "subject"}88 if missing := required.difference(data.columns):89 raise ValueError(f"missing columns: {sorted(missing)}")90 if data[list(required)].isna().any().any():91 raise ValueError("missing model values require an explicit policy")9293 priors = {94 "days": bmb.Prior("Normal", mu=0, sigma=5),95 "1|subject": bmb.Prior("Normal", mu=0, sigma=bmb.Prior("HalfNormal", sigma=10)),96 }97 model = bmb.Model(98 "reaction ~ 1 + days + (1 + days | subject)",99 data,100 family="gaussian",101 priors=priors,102 dropna=False,103 )104 model.build()105 prior = model.prior_predictive(random_seed=seed)106 idata = model.fit(chains=4, random_seed=seed)107 summary = az.summary(idata)108 return model, prior, idata, summary109```110111Prior names and signatures drift; inspect the built model in the installed112Bambi version. The chosen numeric scales are examples, not domain defaults.113114## Non-negotiable rules115116- Do not let Bambi drop rows silently. Reject missing model columns or opt into117 listwise deletion and report the exact retained observation set.118- Formula syntax does not validate causal identification. Confounding,119 post-treatment variables, selection, and measurement error require a model120 justified outside the formula parser.121- Inspect intercept inclusion, categorical reference levels, interactions, and122 group terms. Coefficient meaning changes with coding and predictor centering.123- Automatic prior scaling is a convenience, not a domain prior. Record whether124 `auto_scale` and predictor centering are active and inspect resulting priors.125- A group-specific term requires a meaningful population of groups. Do not use126 fixed dummy effects merely to avoid partial pooling, and do not add random127 slopes unsupported by data without diagnosing geometry.128- `.fit()` returning is not acceptance. Inspect divergences, R-hat, ESS, MCSE,129 traces, prior/posterior predictive checks, and parameterization.130- Predictions on new data must reuse training transforms and category coding.131 State whether they include observation noise and group-specific effects.132- Model comparison requires compatible observations and pointwise log133 likelihood; use ArviZ and inspect Pareto diagnostics.134135Use [testing and version grounding](references/testing-version.md) to inspect136the installed formula, fit, predict, and backend contracts.137138## Completion gate139140Do not declare completion until the formula matches the estimand; response141family/link and auxiliary parameters match support; retained rows and category/142group encodings are explicit; built terms and priors are inspected; prior143predictive and simulated recovery checks pass; inference has no unresolved144critical diagnostics; posterior predictive checks target domain behavior; new-145data prediction semantics and unseen groups are tested; `InferenceData` groups146and coordinates are correct; and skipped checks or version uncertainty are147reported.148149## References150151- [Formula and model contract](references/object-model.md)152- [Families, priors, and inference](references/inference-priors.md)153- [Testing and version grounding](references/testing-version.md)