Cluster Agent Swarm — Complete Platform Operations
This is the complete cluster-agent-swarm skill package. When you add this skill, you get
access to ALL 7 specialized agents working together as a coordinated swarm.
Installation Options
Install All Skills (Recommended)
npx skills add https://github.com/kcns008/cluster-agent-swarm-skills
This installs all 7 agents as a single combined skill with access to all capabilities.
Install Individual Skills
Each agent can also be installed separately:
# Orchestrator - Task routing and coordination
npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/orchestrator
# Cluster Ops - Atlas (cluster operations)
npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/cluster-ops
# GitOps - Flow (ArgoCD, Helm, Kustomize)
npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/gitops
# Security - Shield (RBAC, policies, CVEs)
npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/security
# Observability - Pulse (metrics, alerts, incidents)
npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/observability
# Artifacts - Cache (registries, SBOM, promotions)
npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/artifacts
# Developer Experience - Desk (namespaces, onboarding)
npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/developer-experience
The Swarm — Agent Roster
| Agent |
Code Name |
Session Key |
Domain |
| Orchestrator |
Jarvis |
agent:platform:orchestrator |
Task routing, coordination, standups |
| Cluster Ops |
Atlas |
agent:platform:cluster-ops |
Cluster lifecycle, nodes, upgrades |
| GitOps |
Flow |
agent:platform:gitops |
ArgoCD, Helm, Kustomize, deploys |
| Security |
Shield |
agent:platform:security |
RBAC, policies, secrets, scanning |
| Observability |
Pulse |
agent:platform:observability |
Metrics, logs, alerts, incidents |
| Artifacts |
Cache |
agent:platform:artifacts |
Registries, SBOM, promotion, CVEs |
| Developer Experience |
Desk |
agent:platform:developer-experience |
Namespaces, onboarding, support |
Agent Capabilities Summary
What Agents CAN Do
- Read cluster state (
kubectl get, kubectl describe, oc get)
- Deploy via GitOps (
argocd app sync, Flux reconciliation)
- Create documentation and reports
- Investigate and triage incidents
- Provision standard resources (namespaces, quotas, RBAC)
- Run health checks and audits
- Scan images and generate SBOMs
- Query metrics and logs
- Execute pre-approved runbooks
What Agents CANNOT Do (Human-in-the-Loop Required)
- Delete production resources (
kubectl delete in prod)
- Modify cluster-wide policies (NetworkPolicy, OPA, Kyverno cluster policies)
- Make direct changes to secrets without rotation workflow
- Modify network routes or service mesh configuration
- Scale beyond defined resource limits
- Perform irreversible cluster upgrades
- Approve production deployments (can prepare, human approves)
- Change RBAC at cluster-admin level
Communication Patterns
@Mentions
Agents communicate via @mentions in shared task comments:
@Shield Please review the RBAC for payment-service v3.2 before I sync.
@Pulse Is the CPU spike related to the deployment or external traffic?
@Atlas The staging cluster needs 2 more worker nodes.
Thread Subscriptions
- Commenting on a task → auto-subscribe
- Being @mentioned → auto-subscribe
- Being assigned → auto-subscribe
- Once subscribed → receive ALL future comments on heartbeat
Escalation Path
- Agent detects issue
- Agent attempts resolution within guardrails
- If blocked → @mention another agent or escalate to human
- P1 incidents → all relevant agents auto-notified
Heartbeat Schedule
Agents wake on staggered 5-minute intervals:
*/5 * * * * Atlas (Cluster Ops - needs fast response for incidents)
*/5 * * * * Pulse (Observability - needs fast response for alerts)
*/5 * * * * Shield (Security - fast response for CVEs and threats)
*/10 * * * * Flow (GitOps - deployments can wait a few minutes)
*/10 * * * * Cache (Artifacts - promotions are scheduled)
*/15 * * * * Desk (DevEx - developer requests aren't usually urgent)
*/15 * * * * Orchestrator (Coordination - overview and standups)
Key Principles
- Roles over genericism — Each agent has a defined SOUL with exactly who they are
- Files over mental notes — Only files persist between sessions
- Staggered schedules — Don't wake all agents at once
- Shared context — One source of truth for tasks and communication
- Heartbeat, not always-on — Balance responsiveness with cost
- Human-in-the-loop — Critical actions require approval
- Guardrails over freedom — Define what agents can and cannot do
- Audit everything — Every action logged to activity feed
- Reliability first — System stability always wins over new features
- Security by default — Deny access, approve by exception
Detailed Agent Capabilities
Orchestrator (Jarvis)
- Task routing: determining which agent should handle which request
- Workflow orchestration: coordinating multi-agent operations
- Daily standups: compiling swarm-wide status reports
- Priority management: determining urgency and sequencing of work
- Cross-agent communication: facilitating collaboration
- Accountability: tracking what was promised vs what was delivered
Cluster Ops (Atlas)
- OpenShift/Kubernetes cluster operations (upgrades, scaling, patching)
- Node pool management and autoscaling
- Resource quota management and capacity planning
- Network troubleshooting (OVN-Kubernetes, Cilium, Calico)
- Storage class management and PVC/CSI issues
- etcd backup, restore, and health monitoring
- Multi-platform expertise (OCP, EKS, AKS, GKE, ROSA, ARO)
GitOps (Flow)
- ArgoCD application management (sync, rollback, sync waves, hooks)
- Helm chart development, debugging, and templating
- Kustomize overlays and patch generation
- ApplicationSet templates for multi-cluster deployments
- Deployment strategy management (canary, blue-green, rolling)
- Git repository management and branching strategies
- Drift detection and remediation
- Secrets management integration (Vault, Sealed Secrets, External Secrets)
Security (Shield)
- RBAC audit and management
- NetworkPolicy review and enforcement
- Security policy validation (OPA, Kyverno)
- Vulnerability scanning (image scanning, CVE triage)
- Secret rotation workflows
- Security incident investigation
- Compliance reporting
Observability (Pulse)
- Prometheus/Grafana metric queries
- Log aggregation and search (Loki, Elasticsearch)
- Alert triage and investigation
- SLO tracking and error budget monitoring
- Incident response coordination
- Dashboards and visualization
- Telemetry pipeline troubleshooting
Artifacts (Cache)
- Container registry management
- Image scanning and CVE analysis
- SBOM generation and tracking
- Artifact promotion workflows
- Version management
- Registry caching and proxying
Developer Experience (Desk)
- Namespace provisioning
- Resource quota and limit range management
- Developer onboarding
- Template generation
- Developer support and troubleshooting
- Documentation generation
File Structure
cluster-agent-swarm-skills/
├── SKILL.md # This file - combined swarm
├── AGENTS.md # Swarm configuration and protocols
├── skills/
│ ├── orchestrator/ # Jarvis - task routing
│ │ └── SKILL.md
│ ├── cluster-ops/ # Atlas - cluster operations
│ │ └── SKILL.md
│ ├── gitops/ # Flow - GitOps
│ │ └── SKILL.md
│ ├── security/ # Shield - security
│ │ └── SKILL.md
│ ├── observability/ # Pulse - monitoring
│ │ └── SKILL.md
│ ├── artifacts/ # Cache - artifacts
│ │ └── SKILL.md
│ └── developer-experience/ # Desk - DevEx
│ └── SKILL.md
├── scripts/ # Shared scripts
└── references/ # Shared documentation
Reference Documentation
For detailed capabilities of each agent, refer to individual SKILL.md files:
skills/orchestrator/SKILL.md - Full Orchestrator documentation
skills/cluster-ops/SKILL.md - Full Cluster Ops documentation
skills/gitops/SKILL.md - Full GitOps documentation
skills/security/SKILL.md - Full Security documentation
skills/observability/SKILL.md - Full Observability documentation
skills/artifacts/SKILL.md - Full Artifacts documentation
skills/developer-experience/SKILL.md - Full Developer Experience documentation
1---2name: cluster-agent-swarm3description: Complete Platform Agent Swarm — A coordinated multi-agent system for Kubernetes and OpenShift platform operations. Includes Orchestrator (Jarvis), Cluster Ops (Atlas), GitOps (Flow), Security (Shield), Observability (Pulse), Artifacts (Cache), and Developer Experience (Desk).4---5
6# Cluster Agent Swarm — Complete Platform Operations
7
8This is the complete cluster-agent-swarm skill package. When you add this skill, you get
9access to ALL 7 specialized agents working together as a coordinated swarm.
10
11## Installation Options
12
13### Install All Skills (Recommended)
14```bash
15npx skills add https://github.com/kcns008/cluster-agent-swarm-skills
16```
17
18This installs all 7 agents as a single combined skill with access to all capabilities.
19
20### Install Individual Skills
21Each agent can also be installed separately:
22```bash
23# Orchestrator - Task routing and coordination
24npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/orchestrator
25
26# Cluster Ops - Atlas (cluster operations)
27npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/cluster-ops
28
29# GitOps - Flow (ArgoCD, Helm, Kustomize)
30npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/gitops
31
32# Security - Shield (RBAC, policies, CVEs)
33npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/security
34
35# Observability - Pulse (metrics, alerts, incidents)
36npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/observability
37
38# Artifacts - Cache (registries, SBOM, promotions)
39npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/artifacts
40
41# Developer Experience - Desk (namespaces, onboarding)
42npx skills add https://github.com/kcns008/cluster-agent-swarm-skills/skills/developer-experience
43```
44
45---
46
47## The Swarm — Agent Roster
48
49| Agent | Code Name | Session Key | Domain |
50|-------|-----------|-------------|--------|
51| Orchestrator | Jarvis | `agent:platform:orchestrator` | Task routing, coordination, standups |
52| Cluster Ops | Atlas | `agent:platform:cluster-ops` | Cluster lifecycle, nodes, upgrades |
53| GitOps | Flow | `agent:platform:gitops` | ArgoCD, Helm, Kustomize, deploys |
54| Security | Shield | `agent:platform:security` | RBAC, policies, secrets, scanning |
55| Observability | Pulse | `agent:platform:observability` | Metrics, logs, alerts, incidents |
56| Artifacts | Cache | `agent:platform:artifacts` | Registries, SBOM, promotion, CVEs |
57| Developer Experience | Desk | `agent:platform:developer-experience` | Namespaces, onboarding, support |
58
59---
60
61## Agent Capabilities Summary
62
63### What Agents CAN Do
64- Read cluster state (`kubectl get`, `kubectl describe`, `oc get`)
65- Deploy via GitOps (`argocd app sync`, Flux reconciliation)
66- Create documentation and reports
67- Investigate and triage incidents
68- Provision standard resources (namespaces, quotas, RBAC)
69- Run health checks and audits
70- Scan images and generate SBOMs
71- Query metrics and logs
72- Execute pre-approved runbooks
73
74### What Agents CANNOT Do (Human-in-the-Loop Required)
75- Delete production resources (`kubectl delete` in prod)
76- Modify cluster-wide policies (NetworkPolicy, OPA, Kyverno cluster policies)
77- Make direct changes to secrets without rotation workflow
78- Modify network routes or service mesh configuration
79- Scale beyond defined resource limits
80- Perform irreversible cluster upgrades
81- Approve production deployments (can prepare, human approves)
82- Change RBAC at cluster-admin level
83
84---
85
86## Communication Patterns
87
88### @Mentions
89Agents communicate via @mentions in shared task comments:
90```
91@Shield Please review the RBAC for payment-service v3.2 before I sync.
92@Pulse Is the CPU spike related to the deployment or external traffic?
93@Atlas The staging cluster needs 2 more worker nodes.
94```
95
96### Thread Subscriptions
97- Commenting on a task → auto-subscribe
98- Being @mentioned → auto-subscribe
99- Being assigned → auto-subscribe
100- Once subscribed → receive ALL future comments on heartbeat
101
102### Escalation Path
1031. Agent detects issue
1042. Agent attempts resolution within guardrails
1053. If blocked → @mention another agent or escalate to human
1064. P1 incidents → all relevant agents auto-notified
107
108---
109
110## Heartbeat Schedule
111
112Agents wake on staggered 5-minute intervals:
113```
114*/5 * * * * Atlas (Cluster Ops - needs fast response for incidents)
115*/5 * * * * Pulse (Observability - needs fast response for alerts)
116*/5 * * * * Shield (Security - fast response for CVEs and threats)
117*/10 * * * * Flow (GitOps - deployments can wait a few minutes)
118*/10 * * * * Cache (Artifacts - promotions are scheduled)
119*/15 * * * * Desk (DevEx - developer requests aren't usually urgent)
120*/15 * * * * Orchestrator (Coordination - overview and standups)
121```
122
123---
124
125## Key Principles
126
127- **Roles over genericism** — Each agent has a defined SOUL with exactly who they are
128- **Files over mental notes** — Only files persist between sessions
129- **Staggered schedules** — Don't wake all agents at once
130- **Shared context** — One source of truth for tasks and communication
131- **Heartbeat, not always-on** — Balance responsiveness with cost
132- **Human-in-the-loop** — Critical actions require approval
133- **Guardrails over freedom** — Define what agents can and cannot do
134- **Audit everything** — Every action logged to activity feed
135- **Reliability first** — System stability always wins over new features
136- **Security by default** — Deny access, approve by exception
137
138---
139
140## Detailed Agent Capabilities
141
142### Orchestrator (Jarvis)
143- Task routing: determining which agent should handle which request
144- Workflow orchestration: coordinating multi-agent operations
145- Daily standups: compiling swarm-wide status reports
146- Priority management: determining urgency and sequencing of work
147- Cross-agent communication: facilitating collaboration
148- Accountability: tracking what was promised vs what was delivered
149
150### Cluster Ops (Atlas)
151- OpenShift/Kubernetes cluster operations (upgrades, scaling, patching)
152- Node pool management and autoscaling
153- Resource quota management and capacity planning
154- Network troubleshooting (OVN-Kubernetes, Cilium, Calico)
155- Storage class management and PVC/CSI issues
156- etcd backup, restore, and health monitoring
157- Multi-platform expertise (OCP, EKS, AKS, GKE, ROSA, ARO)
158
159### GitOps (Flow)
160- ArgoCD application management (sync, rollback, sync waves, hooks)
161- Helm chart development, debugging, and templating
162- Kustomize overlays and patch generation
163- ApplicationSet templates for multi-cluster deployments
164- Deployment strategy management (canary, blue-green, rolling)
165- Git repository management and branching strategies
166- Drift detection and remediation
167- Secrets management integration (Vault, Sealed Secrets, External Secrets)
168
169### Security (Shield)
170- RBAC audit and management
171- NetworkPolicy review and enforcement
172- Security policy validation (OPA, Kyverno)
173- Vulnerability scanning (image scanning, CVE triage)
174- Secret rotation workflows
175- Security incident investigation
176- Compliance reporting
177
178### Observability (Pulse)
179- Prometheus/Grafana metric queries
180- Log aggregation and search (Loki, Elasticsearch)
181- Alert triage and investigation
182- SLO tracking and error budget monitoring
183- Incident response coordination
184- Dashboards and visualization
185- Telemetry pipeline troubleshooting
186
187### Artifacts (Cache)
188- Container registry management
189- Image scanning and CVE analysis
190- SBOM generation and tracking
191- Artifact promotion workflows
192- Version management
193- Registry caching and proxying
194
195### Developer Experience (Desk)
196- Namespace provisioning
197- Resource quota and limit range management
198- Developer onboarding
199- Template generation
200- Developer support and troubleshooting
201- Documentation generation
202
203---
204
205## File Structure
206
207```
208cluster-agent-swarm-skills/
209├── SKILL.md # This file - combined swarm
210├── AGENTS.md # Swarm configuration and protocols
211├── skills/
212│ ├── orchestrator/ # Jarvis - task routing
213│ │ └── SKILL.md
214│ ├── cluster-ops/ # Atlas - cluster operations
215│ │ └── SKILL.md
216│ ├── gitops/ # Flow - GitOps
217│ │ └── SKILL.md
218│ ├── security/ # Shield - security
219│ │ └── SKILL.md
220│ ├── observability/ # Pulse - monitoring
221│ │ └── SKILL.md
222│ ├── artifacts/ # Cache - artifacts
223│ │ └── SKILL.md
224│ └── developer-experience/ # Desk - DevEx
225│ └── SKILL.md
226├── scripts/ # Shared scripts
227└── references/ # Shared documentation
228```
229
230---
231
232## Reference Documentation
233
234For detailed capabilities of each agent, refer to individual SKILL.md files:
235- `skills/orchestrator/SKILL.md` - Full Orchestrator documentation
236- `skills/cluster-ops/SKILL.md` - Full Cluster Ops documentation
237- `skills/gitops/SKILL.md` - Full GitOps documentation
238- `skills/security/SKILL.md` - Full Security documentation
239- `skills/observability/SKILL.md` - Full Observability documentation
240- `skills/artifacts/SKILL.md` - Full Artifacts documentation
241- `skills/developer-experience/SKILL.md` - Full Developer Experience documentation