DORA Metrics
Purpose
Compute the four DORA delivery-performance metrics (Deployment
Frequency, Lead Time for Changes, Change Failure Rate, and Time to
Restore Service) from local git history and the GitHub API. Classify
each metric into Elite, High, Medium, or Low using thresholds from
DORA's State of DevOps research, and surface the single weakest
dimension as the next improvement target.
When To Use
- Engineering management retrospectives and quarterly reviews.
- Auditing whether agentic workflows (AI-assisted PRs, automated
deploys) improve velocity and stability or quietly regress them.
- Feeding a tier signal into
minister:release-health-gates.
When Not to Use
- Single-team velocity tracking that needs story-point burndowns
rather than delivery-performance evidence.
- Repositories without a clear production branch or release cadence;
DORA assumes one.
Workflow
Run the helper script with the desired window:
python3 -m minister.dora_metrics --window 30 --branch main
Read the output: per-metric value, tier classification, and the
bottleneck pointer.
For agentic-workflow audits, run the same window twice. Once
filtering to AI-authored PRs (e.g., --failure-label ai-bug),
once across all PRs. Compare the CFR delta. See
modules/agentic-workflow-signals.md.
Optionally pipe --json into the tracker so trend data persists
alongside release-health-gates snapshots.
Optionally render trend charts with kuva when reviewing multiple
windows or comparing before/after an agentic-workflow change:
# Collect weekly snapshots into a TSV, then plot all four metrics
# week<TAB>metric<TAB>value
kuva line trends.tsv --x week --y value --color-by metric \
--title "DORA trends (30-day windows)" -o dora-trends.svg
# Quick terminal preview without writing a file
kuva line trends.tsv --x week --y value --color-by metric --terminal
kuva reads TSV/CSV from stdin or a file path. Install once:
cargo install kuva --features cli. No project source changes
required. See kuva for the
full plot-type reference.
Inputs
| Flag |
Default |
Meaning |
--window |
30 |
Measurement window in days |
--branch |
HEAD |
Production branch |
--failure-label |
bug |
GitHub label marking prod failures |
--json |
off |
Emit JSON instead of human-readable |
--repo-path |
cwd |
Repository directory |
Outputs
A short text report or JSON payload with:
- Per-metric numeric value (e.g.,
4.2/day, 2.1 hours, 8%).
- Per-metric tier (Elite, High, Medium, Low).
- Overall tier (the weakest of the four).
- Bottleneck key, identifying which metric to focus improvement on.
Tier Thresholds
See modules/thresholds.md for the complete table. Brief summary:
| Metric |
Elite |
High |
Medium |
Low |
| DF |
>= 1/day |
>= 1/week |
>= 1/month |
< 1/month |
| LT |
<= 1 day |
<= 1 week |
<= 1 month |
> 1 month |
| CFR |
<= 15% |
<= 30% |
<= 45% |
> 45% |
| TRS |
< 1 hour |
< 1 day |
< 1 week |
>= 1 week |
Verification
Confirm a DORA report is real by re-running the script over a
narrower window and checking that DF and LT scale predictably. For
CFR and TRS, sample two or three of the contributing GitHub issues
and verify the bug (or chosen) label is correct on each.
Testing
Unit tests live in
plugins/minister/tests/unit/test_dora_metrics.py. Each tier
boundary is exercised at the threshold, so future contributors who
adjust an inequality (> vs >=) trigger a failure rather than a
silent regression. Add new tests at the threshold when extending
classification logic.
Exit Criteria
1---2name: dora-metrics-23description: Computes DORA delivery-performance metrics from git and GitHub API. Use when assessing deployment frequency, lead time, or change failure rate.4---56# DORA Metrics78## Purpose910Compute the four DORA delivery-performance metrics (Deployment11Frequency, Lead Time for Changes, Change Failure Rate, and Time to12Restore Service) from local git history and the GitHub API. Classify13each metric into Elite, High, Medium, or Low using thresholds from14DORA's State of DevOps research, and surface the single weakest15dimension as the next improvement target.1617## When To Use1819- Engineering management retrospectives and quarterly reviews.20- Auditing whether agentic workflows (AI-assisted PRs, automated21 deploys) improve velocity and stability or quietly regress them.22- Feeding a tier signal into `minister:release-health-gates`.2324## When Not to Use2526- Single-team velocity tracking that needs story-point burndowns27 rather than delivery-performance evidence.28- Repositories without a clear production branch or release cadence;29 DORA assumes one.3031## Workflow32331. Run the helper script with the desired window:3435 ```bash36 python3 -m minister.dora_metrics --window 30 --branch main37 ```38392. Read the output: per-metric value, tier classification, and the40 bottleneck pointer.41423. For agentic-workflow audits, run the same window twice. Once43 filtering to AI-authored PRs (e.g., `--failure-label ai-bug`),44 once across all PRs. Compare the CFR delta. See45 `modules/agentic-workflow-signals.md`.46474. Optionally pipe `--json` into the tracker so trend data persists48 alongside `release-health-gates` snapshots.49505. Optionally render trend charts with kuva when reviewing multiple51 windows or comparing before/after an agentic-workflow change:5253 ```bash54 # Collect weekly snapshots into a TSV, then plot all four metrics55 # week<TAB>metric<TAB>value56 kuva line trends.tsv --x week --y value --color-by metric \57 --title "DORA trends (30-day windows)" -o dora-trends.svg5859 # Quick terminal preview without writing a file60 kuva line trends.tsv --x week --y value --color-by metric --terminal61 ```6263 kuva reads TSV/CSV from stdin or a file path. Install once:64 `cargo install kuva --features cli`. No project source changes65 required. See [kuva](https://github.com/Psy-Fer/kuva) for the66 full plot-type reference.6768## Inputs6970| Flag | Default | Meaning |71|------|---------|---------|72| `--window` | 30 | Measurement window in days |73| `--branch` | HEAD | Production branch |74| `--failure-label` | bug | GitHub label marking prod failures |75| `--json` | off | Emit JSON instead of human-readable |76| `--repo-path` | cwd | Repository directory |7778## Outputs7980A short text report or JSON payload with:8182- Per-metric numeric value (e.g., `4.2/day`, `2.1 hours`, `8%`).83- Per-metric tier (Elite, High, Medium, Low).84- Overall tier (the weakest of the four).85- Bottleneck key, identifying which metric to focus improvement on.8687## Tier Thresholds8889See `modules/thresholds.md` for the complete table. Brief summary:9091| Metric | Elite | High | Medium | Low |92|--------|-------|------|--------|-----|93| DF | >= 1/day | >= 1/week | >= 1/month | < 1/month |94| LT | <= 1 day | <= 1 week | <= 1 month | > 1 month |95| CFR | <= 15% | <= 30% | <= 45% | > 45% |96| TRS | < 1 hour | < 1 day | < 1 week | >= 1 week |9798## Verification99100Confirm a DORA report is real by re-running the script over a101narrower window and checking that DF and LT scale predictably. For102CFR and TRS, sample two or three of the contributing GitHub issues103and verify the `bug` (or chosen) label is correct on each.104105## Testing106107Unit tests live in108`plugins/minister/tests/unit/test_dora_metrics.py`. Each tier109boundary is exercised at the threshold, so future contributors who110adjust an inequality (`>` vs `>=`) trigger a failure rather than a111silent regression. Add new tests at the threshold when extending112classification logic.113114## Exit Criteria115116- [ ] DORA report generated for the requested window.117- [ ] All four metrics classified into a tier.118- [ ] Bottleneck dimension surfaced.119- [ ] Output is readable in a terminal or as a PR comment.