Ceph AIops
Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by the Ceph project or any storage vendor. Product and trademark names belong to their owners. Source at github.com/AIops-tools/Ceph-AIops under the MIT license.
Governed Ceph operations via the ceph-mgr Dashboard REST API — 37 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.ceph-aiops/, token/runaway budget guard, undo-token recording, and descriptive risk tiers. The Dashboard password is stored encrypted (~/.ceph-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk. The flagship cluster_health turns raw HEALTH_WARN/ERR check codes into plain-language cause + suggested action.
Standalone: the governance harness is bundled in the package (
ceph_aiops.governance) — ceph-aiops has no external skill-family dependency. Works against vanilla ceph-mgr (cephadm / hypervisor-bundled / MicroCeph); no croit, no Kubernetes.
What This Skill Does
| Group | Tools | Count | Read or Write |
|---|---|---|---|
| Health | cluster_health (flagship RCA), cluster_status | 2 | 2 read |
| OSD | osd_tree, osd_df, osd_perf | 3 | 3 read |
| cluster_flag_set, osd_reweight, osd_mark_in, osd_mark_out, osd_purge | 5 | 5 write | |
| PG | pg_summary, pg_dump_stuck, scrub_status | 3 | 3 read |
| trigger_scrub, trigger_deep_scrub | 2 | 2 write | |
| Pool | pool_ls, pool_df | 2 | 2 read |
| set_pool_quota, set_pool_pg_num, set_pool_autoscale, pool_create, set_pool_size, pool_delete | 6 | 6 write | |
| RBD | rbd_ls | 1 | 1 read |
| rbd_image_create, rbd_snapshot_create, rbd_image_delete, rbd_snapshot_delete | 4 | 4 write | |
| CephFS / RGW | cephfs_status, rgw_status | 2 | 2 read |
| Cluster-ops | mon_status, mgr_status, slow_ops, capacity_forecast | 4 | 4 read |
| throttle_recovery | 1 | 1 write | |
| Undo | undo_list, undo_apply | 2 | 2 undo |
Totals: 37 tools — 17 read, 18 write, 2 undo. The MCP server exposes all 37; the CLI is a convenience subset.
Quick Install
uv tool install ceph-aiops
ceph-aiops init # interactive wizard: mgr host/port/username + encrypted Dashboard password
ceph-aiops doctor
When to Use This Skill
- Decode a HEALTH_WARN/ERR state (
cluster_health/health detail) — cause + action per active check (PG_DEGRADED,OSD_NEARFULL,SLOW_OPS,MON_DOWN,LARGE_OMAP_OBJECTS, …) - One-shot triage (
overview): HEALTH status + active checks + OSD up/in counts - Inspect OSDs (
osd_tree/osd_dfmost-full first /osd_perfslowest first), PGs (pg_summary/pg_dump_stuck/scrub_status), pools (pool_ls/pool_dfusable capacity) - Investigate slow requests (
slow_ops), MDS trimming lag (cephfs_status), RGW large-omap (rgw_status), mon quorum (mon_status), and days-to-nearfull (capacity_forecast) - Safely drain + purge an OSD, change pool size/quota, or throttle a slow rebalance (governed writes with dry-run + undo)
Do NOT use when the target is not Ceph (a hypervisor, another storage appliance, a backup product, a container cluster, or a network device). Route those to the appropriate other AIops-tools skill.
Related Skills — Skill Routing
| If the user wants… | Use |
|---|---|
| Ceph: HEALTH_WARN RCA, OSD/PG/pool/RBD/CephFS/RGW, rebalance, slow ops | ceph-aiops (this skill) |
| Any non-Ceph target (hypervisor, other storage, backup, cluster, network) | the appropriate other AIops-tools skill |
Common Workflows
1. "The cluster went HEALTH_WARN overnight" — decode it (read-only)
ceph-aiops doctor→ confirm the mgr Dashboard is reachable and the JWT login works before trusting anything elseceph-aiops overview→ HEALTH status, the list of active check codes, and OSD up/in counts in one shotceph-aiops health detail(MCP:cluster_health) → each active check translated into what it means, the likely cause, and a suggested action- Drill into the implicated resource:
PG_DEGRADED→pg_dump_stuck;OSD_NEARFULL→ceph-aiops osd df(most-full first);SLOW_OPS→slow_ops;MON_DOWN→mon_status;LARGE_OMAP_OBJECTS→rgw_status - Failure branch: if
doctorfails on auth, the Dashboard password is wrong or the store is locked — re-runceph-aiops secret set <target>(or exportCEPH_AIOPS_MASTER_PASSWORDfor non-interactive use). Ifdoctorfails on reachability, the mgrdashboardmodule is likely not enabled; no read is issued against an unauthenticated session.
2. Retire a failing OSD: drain, mark out, purge (governed)
ceph-aiops health detail→ confirm the OSD is genuinely the problem (e.g.OSD_SLOW_PING_TIME, repeatedSLOW_OPSon one id) rather than a cluster-wide symptomceph-aiops osd df→ confirm the id, and that the remaining OSDs have room to absorb its data before you drain anythingceph-aiops osd reweight <id> 0.0→ start a gradual drain; reversible, the prior CRUSH weight is captured as the undo descriptorceph-aiops osd out <id> --dry-run, then re-run without--dry-run→ high risk, double confirmation, needsCEPH_AUDIT_APPROVED_BY- Wait for
ceph-aiops health status/pg_summaryto show all PGsactive+clean— do not purge while backfill is running ceph-aiops osd purge <id> --dry-run, then re-run without--dry-run→ high, irreversible- Failure branch: if client I/O tanks during the drain, stop and reverse —
ceph-aiops undo listthenceph-aiops undo apply <id>restores the prior weight (andosd_mark_inreverses the mark-out). Purge has no undo, which is exactly why it comes last and afteractive+clean.
3. Recovery is starving client I/O
ceph-aiops health detail→ confirm the cluster is actually backfilling/recovering (PG_DEGRADED,PG_BACKFILL_FULL) rather than hitting a different bottleneckslow_ops→ check whether client requests are genuinely being blocked, and by which OSDsthrottle_recovery(max_backfills=1, recovery_max_active=1)→ med risk, reversible; the priorosd_max_backfills/osd_recovery_max_activeare captured as the undo descriptor- Re-check
slow_opsandpg_summary— recovery is slower but client latency should recover - Once the cluster is quiet, raise the values back (or
ceph-aiops undo apply <id>to restore the exact prior settings) - Failure branch: if throttling does not help, the bottleneck is not recovery — go back to
osd_perf(slowest OSDs first) andmon_status; do not keep lowering the throttle, you will only extend the degraded window.
4. A pool is running out of usable capacity
ceph-aiops overview→ look forPOOL_NEARFULL/OSD_NEARFULLamong the active checkspool_df→ per-pool usage with usable capacity = raw ÷ size (asize=3pool reports a third of raw — this is where most "but the disks aren't full" confusion comes from)capacity_forecast→ days-to-nearfull at the current fill rate, so you know whether this is a this-week problem or a this-quarter one- Buy time reversibly first:
set_pool_quota(med, undo → prior quota) orset_pool_autoscale(med, undo) to let pg_num track the new size - Only if a replica change is genuinely the answer:
set_pool_size --dry-runthen the real call — high risk, because loweringsizereduces durability and any change forces cluster-wide data movement - Failure branch: if the resulting rebalance saturates the cluster, apply workflow 3 (
throttle_recovery) rather than reverting the size mid-flight; if the size change itself was wrong,ceph-aiops undo apply <id>replays the recorded prior value — but expect a second full rebalance.
Governance & Safety
The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (a ceph-mgr Dashboard account with a read-only role — writes then fail at the mgr). There is no read-only switch, policy file, or approval gate.
- Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to
~/.ceph-aiops/audit.db(relocatable viaCEPH_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does. CEPH_AUDIT_APPROVED_BY/CEPH_AUDIT_RATIONALEare optional annotations recorded on the audit row (who/why); they are never required and never block.- Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with
CEPH_RUNAWAY_MAX=0. - Destructive writes support
--dry-run/dry_run=Trueand double confirmation at the CLI. - Reversible writes fetch the real before-state and record an inverse descriptor (
osd_reweight→restore prior weight,cluster_flag_set→toggle back); irreversible ops (osd_purge,pool_delete, RBD deletes) record only the before-state.
References
references/capabilities.md— full tool → API-path → returns referencereferences/cli-reference.md— CLI command referencereferences/setup-guide.md— onboarding, credentials, and connectivity