🔧 DevOps Challenge
Master the deployment and operations of AI agent infrastructure! This challenge is for DevOps and Platform engineers who want to ensure Aden runs reliably at scale.
Difficulty: Advanced Time: 2-3 hours Prerequisites: Complete Getting Started, Docker, Linux, CI/CD experience
Part 1: Infrastructure Analysis (20 points)
Task 1.1: Docker Deep Dive 🐳
Analyze the Aden Docker setup:
- What Dockerfile exists in the repository and what does it build?
- How would you containerize the MCP tools server?
- How is hot reload enabled for development?
- What would need to be mounted as volumes for persistence?
- What networking considerations exist for the MCP server?
Task 1.2: Service Dependencies 🔗
Map the service dependencies:
- Create a dependency diagram showing which services depend on which
- What's the startup order? Does it matter?
- What happens if MongoDB is unavailable?
- What happens if Redis is unavailable?
- Which services are stateless vs stateful?
Task 1.3: Configuration Management ⚙️
Analyze how configuration works:
- How does
config.yamlget generated? - What environment variables are required?
- How are secrets managed? (API keys, database passwords)
- What's the difference between dev and prod configs?
Part 2: Deployment Scenarios (25 points)
Task 2.1: Production Deployment Plan 📋
Design a production deployment for a company with:
- 100 active agents
- 10,000 LLM requests/day
- 99.9% uptime requirement
- Multi-region support needed
Provide:
- Infrastructure diagram (cloud provider of your choice)
- Service sizing (CPU, memory for each component)
- Database setup (primary/replica, backups)
- Load balancing strategy
- Estimated monthly cost
Task 2.2: Kubernetes Migration 🚢
Convert the Docker Compose setup to Kubernetes:
- Create a Kubernetes deployment manifest for the Hive backend
- Create a Service and Ingress for external access
- Design a ConfigMap for configuration
- Create a Secret for sensitive data
- Set up a HorizontalPodAutoscaler
# Provide your manifests here
apiVersion: apps/v1
kind: Deployment
metadata:
name: hive-backend
spec:
# Your implementation
Task 2.3: High Availability Design 🔄
Design for high availability:
- How would you handle backend service failures?
- How would you handle database failover?
- What's your strategy for zero-downtime deployments?
- How would you handle WebSocket connections during rolling updates?
- Design a disaster recovery plan
Part 3: CI/CD Pipeline (25 points)
Task 3.1: GitHub Actions Pipeline 🔄
Create a complete CI/CD pipeline:
# .github/workflows/ci-cd.yml
name: Aden CI/CD
on:
push:
branches: [main, develop]
pull_request:
branches: [main]
jobs:
# Your implementation should include:
# - Linting
# - Type checking
# - Unit tests
# - Integration tests
# - Build Docker images
# - Push to registry
# - Deploy to staging (on develop)
# - Deploy to production (on main, with approval)
Include:
- Separate jobs for frontend and backend
- Matrix testing for multiple Node versions
- Docker layer caching
- Deployment gates/approvals
- Rollback strategy
Task 3.2: Testing Strategy 🧪
Design the testing infrastructure:
- Unit Tests: What to test? How to mock LLM calls?
- Integration Tests: How to test with real databases?
- E2E Tests: What user flows to test?
- Load Tests: How to simulate agent traffic?
- Chaos Tests: What failures to simulate?
Provide example test configurations for each type.
Task 3.3: Environment Management 🌍
Design environment strategy:
| Environment | Purpose | Data | Who Can Access |
|---|---|---|---|
| Local | Development | Mock | Developers |
| Dev | Integration | Sanitized | Engineering |
| Staging | Pre-prod | Copy of prod | Engineering + QA |
| Production | Live | Real | Restricted |
For each environment, specify:
- How it's provisioned
- How data is managed
- How deployments happen
- Access control
Part 4: Observability & Operations (30 points)
Task 4.1: Monitoring Stack 📊
Design a comprehensive monitoring solution:
- Metrics: What to collect? (list at least 10 key metrics)
- Logs: Logging strategy and aggregation
- Traces: Distributed tracing for agent flows
- Dashboards: Design 3 key dashboards
# Provide a docker-compose addition for monitoring
services:
prometheus:
# Your config
grafana:
# Your config
# Add more as needed
Task 4.2: Alerting Rules 🚨
Create alerting rules for critical scenarios:
# Prometheus alerting rules
groups:
- name: aden-critical
rules:
- alert: HighErrorRate
expr: # Your expression
for: 5m
labels:
severity: critical
annotations:
summary: "High error rate detected"
description: # Your description
# Add more alerts for:
# - Service down
# - High latency
# - Budget exceeded
# - Database connection issues
# - Memory pressure
Create at least 8 alert rules covering different failure modes.
Task 4.3: Incident Response 🆘
Create an incident response runbook:
Scenario: Agent response times spike to 30 seconds (normal: 2 seconds)
Provide:
- Detection: How was this discovered?
- Triage: Initial investigation steps
- Diagnosis: Decision tree for root causes
- Resolution: Steps for each root cause
- Post-mortem: Template for incident review
# Runbook: High Agent Latency
## Symptoms
- Agent response times > 10s
- Dashboard showing degraded status
## Initial Triage
1. Check [ ] Is this affecting all agents or specific ones?
2. Check [ ] Is the backend healthy? (health endpoint)
3. Check [ ] Are databases responsive?
...
## Diagnostic Steps
...
## Resolution Steps
### If LLM Provider Issue:
...
### If Database Issue:
...
Part 5: Security Hardening (Bonus - 20 points)
Task 5.1: Security Audit 🔒
Perform a security analysis:
- Network: What ports are exposed? Are they necessary?
- Secrets: How are secrets currently handled? Improvements?
- Authentication: How is API auth implemented?
- Container Security: What image scanning would you add?
- Database Security: What hardening is needed?
Task 5.2: Compliance Checklist ✅
For SOC 2 compliance, what changes are needed?
- Access control improvements
- Audit logging requirements
- Encryption requirements
- Data retention policies
- Incident response requirements
Submission Checklist
- Part 1 infrastructure analysis
- Part 2 deployment designs and manifests
- Part 3 CI/CD pipeline YAML
- Part 4 monitoring and alerting configs
- (Bonus) Part 5 security analysis
How to Submit
- Create a GitHub Gist with your answers
- Name it
aden-devops-YOURNAME.md - Include all YAML/configuration files
- Include any diagrams (use Mermaid, ASCII, or image links)
- Email to
careers@adenhq.com- Subject:
[DevOps Challenge] Your Name
- Subject:
Scoring
| Section | Points |
|---|---|
| Part 1: Infrastructure | 20 |
| Part 2: Deployment | 25 |
| Part 3: CI/CD | 25 |
| Part 4: Observability | 30 |
| Part 5: Security (Bonus) | +20 |
| Total | 100 (+20) |
Passing score: 75+ points
Bonus Points (+15)
- +5: Set up a working local Kubernetes cluster with Aden
- +5: Create a Terraform module for cloud deployment
- +5: Submit a PR improving deployment documentation
Resources
Good luck! We're looking for engineers who keep systems running smoothly! 🔧✨