Inference AIops
Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by the vLLM or Ray projects or any inference-serving vendor. Product and trademark names belong to their owners. Source at github.com/AIops-tools/Inference-AIops under the MIT license.
Governed GPU-inference operations for vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI — 39 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.inference-aiops/, policy engine, token/runaway budget guard, undo-token recording, and descriptive risk-tier labels on every audit row. The flagship diagnose_latency_spike folds queue depth + KV-cache pressure + prefix-cache locality into a ranked cause and the specific knob to turn; the engine-agnostic diagnose_engine_latency does the same across whatever signals SGLang/TGI expose. Each engine's Prometheus /metrics is parsed directly — no Prometheus server required.
Standalone: the governance harness is bundled in the package (
inference_aiops.governance) — no external skill-family dependency. A bearer token is optional (many stacks run open).
What This Skill Does
| Group | Tools | Count | Read or Write |
|---|---|---|---|
| Metrics & RCA (vLLM) | request metrics, queue depth, KV-cache stats, diagnose latency spike, diagnose low utilisation | 5 | 5 read |
| Engine-agnostic (vLLM/SGLang/TGI) | engine health, engine inventory, engine request metrics, engine queue depth, diagnose engine latency | 5 | 5 read |
| Ray Serve (read) | deployment list, deployment status, replica list, autoscale config get | 4 | 4 read |
| Ray Serve (write) | scale up (med), scale down (high), scale-to-zero (high), autoscale config update (med), drain replica (high) | 5 | 5 write |
| Models / vLLM | model list, model info, LoRA load (med), LoRA unload (high), base hot-swap (high) | 5 | 2 read / 3 write |
| Ray cluster / jobs / GPU | cluster resources, dashboard status, job list, GPU utilisation, job cancel (med), replica restart (high) | 6 | 4 read / 2 write |
| Deploy lifecycle | deploy (med), undeploy (high), redeploy (high), routing policy update (med) | 4 | 4 write |
| Cost | cost per token | 1 | 1 read |
23 read, 16 write, plus undo_list / undo_apply — 39 MCP tools in total. The high-risk writes support dry_run + double-confirm; reversible writes record an undo descriptor. The engine-agnostic reads cover any engine; the Ray Serve / cluster / deploy write groups are vLLM-only and teach-and-refuse on a SGLang/TGI target (single-process engines have no Ray control plane).
Quick Install
uv tool install inference-aiops
inference-aiops init # interactive wizard: engine (vllm/sglang/tgi) + host + port + scheme (token optional)
inference-aiops doctor # vLLM: probes Ray + vLLM; SGLang/TGI: engine health + inventory
When to Use This Skill
- Triage a cluster (
overview): Serve deployments, total replicas, queue backpressure - Diagnose slow inference (
metrics diagnose/diagnose_latency_spike): rank the cause (queue depth vs KV-cache preemption vs prefix-cache locality) and get the knob to turn - Find idle GPUs and over-provisioned replicas (
diagnose_low_utilization) - Scale a Ray Serve deployment up/down, scale-to-zero to stop cost bleed, or update autoscale bounds
- Drain a replica gracefully before a node reboot (finishes in-flight requests)
- Load/unload a LoRA adapter; hot-swap a base model (Sleep-Mode swap, captures the prior model)
- Inspect GPU utilisation per node, list/cancel Ray jobs, restart a stuck replica
- Compute cost per million tokens from throughput × GPU $/hr
- Observe an SGLang or TGI server (
engine_health,engine_inventory,engine_request_metrics,engine_queue_depth,diagnose_engine_latency) — single-process engines with no Ray control plane
Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools. This skill is scoped to GPU inference serving (vLLM + Ray).
Related Skills — Skill Routing
| If the user wants… | Use |
|---|---|
| vLLM / Ray Serve inference: latency RCA, autoscale, drain, LoRA, cost/token | inference-aiops (this skill) |
| SGLang / TGI serving: health, running-model inventory, request metrics, queue depth, latency RCA | inference-aiops (this skill — engine-agnostic reads) |
| Any non-inference infrastructure (hypervisor, storage, backup, general clusters, network, OT) | the appropriate other AIops-tools line |
Common Workflows
1. "Inference got slow this afternoon" (flagship RCA → the right knob)
inference-aiops doctor→ confirm the vLLM endpoint and Ray dashboard are actually reachable before blaming the modelinference-aiops overview→ Serve deployments, total replicas, and whether queue backpressure is cluster-wide or one deploymentinference-aiops metrics diagnose(MCP:diagnose_latency_spike) → a ranked cause with the measured numbers: iswaitingqueue depth high (backpressure)? Are there KV-cache preemptions (kv_cache_stats)? Has the prefix-cache hit rate dropped (routing lost locality)?- Turn the knob the RCA names, not a guess:
- backpressure →
inference-aiops serve scale <app> <deployment> --replicas N(scale_replicas_up, reversible, prior count captured) - KV-cache preemption →
autoscale_config_updateto lower the concurrent-request cap (reversible, prior config captured) - lost locality →
routing_policy_updateto prefix-aware / session-affinity (reversible)
- backpressure →
- Re-check
inference-aiops metrics requests(TTFT / TPOT / e2e) andinference-aiops metrics queueto confirm the p99 actually moved - Failure branch: if the fix makes it worse,
inference-aiops undo list→inference-aiops undo apply <id>restores the exact prior replica count / autoscale config / routing policy. Ifdiagnose_latency_spikereports no clear cause, the bottleneck is likely upstream of serving — checkgpu_utilizationfor a throttling or shared-GPU problem before scaling anything.
2. Off-peak cost save: scale a deployment down to zero and bring it back
inference-aiops metrics requests→ confirm traffic really is idle, not just briefly quietdiagnose_low_utilization→ the deployments actually burning GPU for nothing, with the measured utilisationcost_per_token→ quantify the bleed ($/1M tokens at the current throughput) so the change is justifiable in the audit trail- (optional)
export INFERENCE_AUDIT_APPROVED_BY=you INFERENCE_AUDIT_RATIONALE="off-peak cost save"→ annotates the audit row with who/why; recorded when set, never required inference-aiops serve scale-to-zero <app> <deployment> --dry-run, then re-run without--dry-run→ high risk, double confirmation.scale_to_zerostops the bleed but strands ingress — requests will queue or fail until replicas return- To restore:
inference-aiops undo apply <id>(replays the captured prior replica count) orinference-aiops serve scale <app> <deployment> --replicas N - Failure branch: if traffic arrives while at zero, restore immediately via undo — do not wait for autoscale, since
scale_to_zeromay have been applied outside the autoscaler's floor. If the restore fails,serve statuswill show the deployment unhealthy;deployment_redeployis the last resort (high risk, disruptive).
3. Drain a replica before a node reboot
inference-aiops serve list/replica_list→ identify the replicas pinned to the node you are about to rebootqueue_depth→ confirm the remaining replicas can absorb the load; if not,scale_replicas_upfirst so draining does not cause a brownoutdrain_replica <app> <deployment> <replica_id> --dry-run, then confirm → high risk; the drain finishes in-flight requests before removing the replica- Watch
replica_listuntil the replica is gone andrequest_metricsshows no error spike, then reboot the node - Failure branch: if the drain hangs on a long-running request,
replica_restartforcibly cycles it — that drops in-flight requests, so only reach for it once you accept the loss. Multi-node drain has not been verified against a live cluster (seedocs/VERIFICATION.md).
4. Free GPU memory between bursts with Sleep Mode, then resume
model_is_sleeping→ is the engine already suspended?nullmeans the engine did not report it — that is UNKNOWN, not awake, so resolve it before writingrequest_metrics/queue_depth→ confirm the engine is actually idle; sleeping a busy engine drops live trafficmodel_sleep --dry-run, then confirm → high risk. Level 1 offloads the weights to CPU RAM and wakes fast; level 2 discards them, so waking reloads from disk. The undo descriptor is recorded only if the engine was observed awake first — an already-sleeping engine records none, so an undo can never wake something this call did not suspend- Verify:
model_is_sleepingreports true, and GPU memory has been released (gpu_utilization) - Resume with
model_wake(medium risk), orinference-aiops undo apply <id>to replay the recorded inverse.model_wakeitself records no undo: vLLM reports whether the engine sleeps but never at which level, and guessing between level 1 and level 2 would be inventing a prior state - Failure branch: if any of the three tools reports that the route does not exist, the server was not started with
VLLM_SERVER_DEV_MODE=1. That is a server start-up flag, not a fault in the tool and not a stale id — restart vLLM with the flag, or leave Sleep Mode off if this is a production deployment that should not expose it.
vLLM has no in-place base-model swap. Sleep Mode suspends and resumes the same model; serving a different base model means restarting vLLM with a different
--model. For adapter-level changes uselora_load(reversible) andlora_unload(high).
Governance & Safety
The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the environment you connect it with (a network path that only reaches the read/metrics endpoints, a Ray dashboard without its job-submission API — writes then fail at the server). There is no read-only switch, policy file, or approval gate.
- Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to
~/.inference-aiops/audit.db(relocatable viaINFERENCE_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does. INFERENCE_AUDIT_APPROVED_BY/INFERENCE_AUDIT_RATIONALEare optional annotations recorded on the audit row (who/why); they are never required and never block.- Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker.
- The fragile prod writes support
--dry-run/dry_run=Trueand double confirmation at the CLI. - Reversible writes (scale, autoscale-config, routing, hot-swap, LoRA load) capture before-state and record an inverse descriptor.
References
references/capabilities.md— full tool → backend → endpoint → returns referencereferences/cli-reference.md— CLI command referencereferences/setup-guide.md— onboarding, optional token, and connectivity