ArviZ Python
Turn posterior artifacts into correctly labeled inference data, preserve chain
and draw structure, diagnose sampling and prediction, and compare only models
whose observations and log-likelihood semantics are compatible.
Boundary
Use this skill when the input is completed or partially completed Bayesian
draws and the code directly uses ArviZ. Use the modeling framework's skill to
build or sample a model. A plot request alone does not trigger this skill unless
the plot represents posterior, diagnostic, predictive, or comparison results.
Know the objects
| Object |
Meaning |
Use it for |
ArviZ 1.x DataTree / accepted idata-like container |
Named inference groups, each holding labeled xarray variables. Legacy integrations may still call the artifact InferenceData. |
The durable boundary for posterior analysis. |
| Group |
A semantic dataset such as posterior, sample_stats, log_likelihood, posterior_predictive, predictions, observed_data, or constant_data. |
Keeping quantities with different meaning separate. |
| Variable |
A named tensor with dimensions and coordinates. |
One parameter, statistic, observation, or prediction. |
chain |
Independent MCMC run dimension. |
Between-chain convergence diagnostics. |
draw |
Iteration within a chain. |
Within-chain sample sequence. |
| Model-comparison result |
Pointwise predictive score summaries, uncertainty, diagnostics, and weights. |
Comparing compatible predictive models, not proving truth. |
Dimension names determine semantics; axis order is secondary. Sample dimensions
must not be confused with event/data dimensions. observed_data has no chain or
draw axis; posterior predictive values normally have both plus observation
dimensions. Read the inference-data contract before
converting arrays, concatenating results, or selecting coordinates.
Ordered workflow
- Establish provenance: model, sampler/guide, package versions, warmup policy,
number of chains/draws, seed, observation identity, and log-likelihood target.
- Inspect available groups, variables, dimensions, coordinates, dtypes, and
missing values before computing a statistic. On ArviZ 1.x, treat the
DataTree children as the group registry instead of assuming 0.x
InferenceData attribute behavior.
- Convert raw arrays with explicit group, variable, dims, and coords mappings.
Preserve chain and draw separately; reject ambiguous flattened samples.
- Select variables/coordinates intentionally. Exclude warmup or transformed
helper variables only under a recorded policy.
- Diagnose sampling using multiple complementary measures and plots. Report
problematic variables and coordinates, not only a global maximum.
- Check posterior predictive adequacy with discrepancies tied to the modeling
objective and distinguish in-sample replicated data from out-of-sample
predictions.
- For model comparison, verify identical observations, likelihood target, data
preprocessing, and pointwise
log_likelihood; inspect Pareto diagnostics and
score uncertainty before ranking.
- Persist the labeled artifact and provenance when results must be reproducible.
Decision map
| Need |
Use |
Guard |
| Describe marginal posterior |
Summary/HDI plus MCSE and ESS |
Do not report precision beyond Monte Carlo accuracy. |
| Diagnose MCMC convergence |
Rank R-hat, bulk/tail ESS, MCSE, trace/rank/energy views, sampler stats |
Keep chains separate and inspect parameter coordinates. |
| Diagnose one-chain output |
ESS/MCSE and trace/autocorrelation evidence |
R-hat cannot supply between-chain evidence; obtain more chains where possible. |
| Assess model fit |
Posterior predictive checks |
Choose domain-relevant discrepancy, not only overlapping histograms. |
| Compare predictive models |
PSIS-LOO or justified alternative |
Same observations/target and valid pointwise log likelihood; inspect Pareto-k. |
| Combine independent chains |
Concatenate along chain only after schema/coordinate equality |
Do not concatenate posterior draws along an observation dimension. |
| Combine different groups |
Extend/merge by group under matching variables/coords |
Do not overwrite an existing group silently. |
Read diagnostics and comparison before
turning any single number into a pass/fail conclusion.
Canonical labeled conversion
import arviz as az
import numpy as np
def make_idata(
beta: np.ndarray,
log_likelihood: np.ndarray,
feature_names: list[str],
observation_ids: list[str],
):
if beta.ndim != 3 or log_likelihood.ndim != 3:
raise ValueError("expected chain x draw x domain arrays")
if beta.shape[:2] != log_likelihood.shape[:2]:
raise ValueError("posterior and log likelihood sample axes differ")
if beta.shape[2] != len(feature_names):
raise ValueError("feature coordinate length differs")
if log_likelihood.shape[2] != len(observation_ids):
raise ValueError("observation coordinate length differs")
if len(set(feature_names)) != len(feature_names):
raise ValueError("feature coordinates must be unique")
if len(set(observation_ids)) != len(observation_ids):
raise ValueError("observation coordinates must be unique")
return az.from_dict(
{
"posterior": {"beta": beta},
"log_likelihood": {"outcome": log_likelihood},
},
coords={"feature": feature_names, "observation": observation_ids},
dims={"beta": ["feature"], "outcome": ["observation"]},
)
This assumes arrays arrive as chain × draw × domain. Inspect the installed
converter rather than applying that assumption to every backend. Assert
coordinate lengths and uniqueness before conversion.
Diagnostic contract
- R-hat near one is necessary for many MCMC workflows but not sufficient. Use
rank-normalized and folded variants as supported, plus trace behavior.
- Bulk ESS concerns central estimates; tail ESS concerns quantiles/tails. MCSE
must be small relative to the precision needed for the reported estimand.
- Divergences, energy/BFMI, acceptance, and tree-depth fields live in
sample_stats when the sampler records them. Absence is unknown, not zero.
- Do not average chains before diagnostics or reshape chain × draw to one axis
and then recreate fake chains.
- Diagnostic thresholds are escalation rules, not proof of model validity.
Predictive adequacy and model assumptions remain separate.
Comparison contract
PSIS-LOO requires pointwise log likelihood for the same observation units. Do
not compare models fit to different filtered rows, likelihood factorizations,
response transformations, or weighting conventions without a justified mapping.
On ArviZ 1.x, call loo(..., pointwise=True, var_name=...) when the caller must
retain per-observation Pareto-k, then pass those ELPD results to compare.
compare no longer accepts the 0.x ic="loo" selector. Inspect Pareto-k and
uncertainty; address influential observations or use a more robust validation
plan rather than reporting weights mechanically. Comparison answers relative
predictive performance among candidates, not absolute fit, causality, or
scientific truth.
Use testing and version grounding because
ArviZ 1.x packaging and stats/plot APIs differ from many 0.x examples.
Completion gate
Do not declare completion until provenance is recorded; required groups exist;
each variable's dimensions and coordinates match semantic axes; chains and
draws remain separate; convergence reports combine R-hat, ESS, MCSE, traces,
and available sampler statistics; predictive checks target the stated use;
comparisons use identical observations and valid pointwise log likelihood;
Pareto and uncertainty warnings are surfaced; persisted output reloads with the
same groups/coords; and missing groups or skipped diagnostics are reported.
References
- Inference-data contract
- Diagnostics and comparison
- Testing and version grounding
1---2name: arviz-python3description: Use for writing, reviewing, debugging, or testing Python analysis of Bayesian inference results with ArviZ, including 1.x DataTree groups, legacy InferenceData inputs, xarray dimensions and coordinates, conversion, summaries, R-hat/ESS/MCSE diagnostics, posterior predictive checks, PSIS-LOO, Pareto-k, and model comparison. Trigger on chain/draw shape errors, mislabeled groups, flattened samples, missing log likelihood, or misleading diagnostic claims. Do not use to construct or sample PyMC, NumPyro, or Bambi models, for generic plotting, or for deterministic statistics without Bayesian draws.4---56# ArviZ Python78Turn posterior artifacts into correctly labeled inference data, preserve chain9and draw structure, diagnose sampling and prediction, and compare only models10whose observations and log-likelihood semantics are compatible.1112## Boundary1314Use this skill when the input is completed or partially completed Bayesian15draws and the code directly uses ArviZ. Use the modeling framework's skill to16build or sample a model. A plot request alone does not trigger this skill unless17the plot represents posterior, diagnostic, predictive, or comparison results.1819## Know the objects2021| Object | Meaning | Use it for |22|---|---|---|23| ArviZ 1.x `DataTree` / accepted idata-like container | Named inference groups, each holding labeled xarray variables. Legacy integrations may still call the artifact `InferenceData`. | The durable boundary for posterior analysis. |24| Group | A semantic dataset such as `posterior`, `sample_stats`, `log_likelihood`, `posterior_predictive`, `predictions`, `observed_data`, or `constant_data`. | Keeping quantities with different meaning separate. |25| Variable | A named tensor with dimensions and coordinates. | One parameter, statistic, observation, or prediction. |26| `chain` | Independent MCMC run dimension. | Between-chain convergence diagnostics. |27| `draw` | Iteration within a chain. | Within-chain sample sequence. |28| Model-comparison result | Pointwise predictive score summaries, uncertainty, diagnostics, and weights. | Comparing compatible predictive models, not proving truth. |2930Dimension names determine semantics; axis order is secondary. Sample dimensions31must not be confused with event/data dimensions. `observed_data` has no chain or32draw axis; posterior predictive values normally have both plus observation33dimensions. Read [the inference-data contract](references/object-model.md) before34converting arrays, concatenating results, or selecting coordinates.3536## Ordered workflow37381. Establish provenance: model, sampler/guide, package versions, warmup policy,39 number of chains/draws, seed, observation identity, and log-likelihood target.402. Inspect available groups, variables, dimensions, coordinates, dtypes, and41 missing values before computing a statistic. On ArviZ 1.x, treat the42 `DataTree` children as the group registry instead of assuming 0.x43 `InferenceData` attribute behavior.443. Convert raw arrays with explicit group, variable, dims, and coords mappings.45 Preserve chain and draw separately; reject ambiguous flattened samples.464. Select variables/coordinates intentionally. Exclude warmup or transformed47 helper variables only under a recorded policy.485. Diagnose sampling using multiple complementary measures and plots. Report49 problematic variables and coordinates, not only a global maximum.506. Check posterior predictive adequacy with discrepancies tied to the modeling51 objective and distinguish in-sample replicated data from out-of-sample52 predictions.537. For model comparison, verify identical observations, likelihood target, data54 preprocessing, and pointwise `log_likelihood`; inspect Pareto diagnostics and55 score uncertainty before ranking.568. Persist the labeled artifact and provenance when results must be reproducible.5758## Decision map5960| Need | Use | Guard |61|---|---|---|62| Describe marginal posterior | Summary/HDI plus MCSE and ESS | Do not report precision beyond Monte Carlo accuracy. |63| Diagnose MCMC convergence | Rank R-hat, bulk/tail ESS, MCSE, trace/rank/energy views, sampler stats | Keep chains separate and inspect parameter coordinates. |64| Diagnose one-chain output | ESS/MCSE and trace/autocorrelation evidence | R-hat cannot supply between-chain evidence; obtain more chains where possible. |65| Assess model fit | Posterior predictive checks | Choose domain-relevant discrepancy, not only overlapping histograms. |66| Compare predictive models | PSIS-LOO or justified alternative | Same observations/target and valid pointwise log likelihood; inspect Pareto-k. |67| Combine independent chains | Concatenate along `chain` only after schema/coordinate equality | Do not concatenate posterior draws along an observation dimension. |68| Combine different groups | Extend/merge by group under matching variables/coords | Do not overwrite an existing group silently. |6970Read [diagnostics and comparison](references/diagnostics-comparison.md) before71turning any single number into a pass/fail conclusion.7273## Canonical labeled conversion7475```python76import arviz as az77import numpy as np787980def make_idata(81 beta: np.ndarray,82 log_likelihood: np.ndarray,83 feature_names: list[str],84 observation_ids: list[str],85):86 if beta.ndim != 3 or log_likelihood.ndim != 3:87 raise ValueError("expected chain x draw x domain arrays")88 if beta.shape[:2] != log_likelihood.shape[:2]:89 raise ValueError("posterior and log likelihood sample axes differ")90 if beta.shape[2] != len(feature_names):91 raise ValueError("feature coordinate length differs")92 if log_likelihood.shape[2] != len(observation_ids):93 raise ValueError("observation coordinate length differs")94 if len(set(feature_names)) != len(feature_names):95 raise ValueError("feature coordinates must be unique")96 if len(set(observation_ids)) != len(observation_ids):97 raise ValueError("observation coordinates must be unique")98 return az.from_dict(99 {100 "posterior": {"beta": beta},101 "log_likelihood": {"outcome": log_likelihood},102 },103 coords={"feature": feature_names, "observation": observation_ids},104 dims={"beta": ["feature"], "outcome": ["observation"]},105 )106```107108This assumes arrays arrive as chain × draw × domain. Inspect the installed109converter rather than applying that assumption to every backend. Assert110coordinate lengths and uniqueness before conversion.111112## Diagnostic contract113114- R-hat near one is necessary for many MCMC workflows but not sufficient. Use115 rank-normalized and folded variants as supported, plus trace behavior.116- Bulk ESS concerns central estimates; tail ESS concerns quantiles/tails. MCSE117 must be small relative to the precision needed for the reported estimand.118- Divergences, energy/BFMI, acceptance, and tree-depth fields live in119 `sample_stats` when the sampler records them. Absence is unknown, not zero.120- Do not average chains before diagnostics or reshape chain × draw to one axis121 and then recreate fake chains.122- Diagnostic thresholds are escalation rules, not proof of model validity.123 Predictive adequacy and model assumptions remain separate.124125## Comparison contract126127PSIS-LOO requires pointwise log likelihood for the same observation units. Do128not compare models fit to different filtered rows, likelihood factorizations,129response transformations, or weighting conventions without a justified mapping.130On ArviZ 1.x, call `loo(..., pointwise=True, var_name=...)` when the caller must131retain per-observation Pareto-k, then pass those ELPD results to `compare`.132`compare` no longer accepts the 0.x `ic="loo"` selector. Inspect Pareto-k and133uncertainty; address influential observations or use a more robust validation134plan rather than reporting weights mechanically. Comparison answers relative135predictive performance among candidates, not absolute fit, causality, or136scientific truth.137138Use [testing and version grounding](references/testing-version.md) because139ArviZ 1.x packaging and stats/plot APIs differ from many 0.x examples.140141## Completion gate142143Do not declare completion until provenance is recorded; required groups exist;144each variable's dimensions and coordinates match semantic axes; chains and145draws remain separate; convergence reports combine R-hat, ESS, MCSE, traces,146and available sampler statistics; predictive checks target the stated use;147comparisons use identical observations and valid pointwise log likelihood;148Pareto and uncertainty warnings are surfaced; persisted output reloads with the149same groups/coords; and missing groups or skipped diagnostics are reported.150151## References152153- [Inference-data contract](references/object-model.md)154- [Diagnostics and comparison](references/diagnostics-comparison.md)155- [Testing and version grounding](references/testing-version.md)